models · qwen

Qwen3.5-122B-A10B

Qwen3.5 MoE 122B total / 10B active. 262K context, 129 tok/s. Arena 750. GPQA Diamond 86.6%. Strong MoE efficiency at near-70B inference cost.

Compare Verify

What this registry has checked

Quarantine. We hold the publisher’s own published hash for these weights and have recorded it against this entry. We have not downloaded and hashed the weights ourselves, so the hash is the publisher’s claim, not our measurement.
Weight file SHA-2569c064915c4742a03e8d5e317e3841f521e9bf51a5cedfe3829d4e173b13555a5
Methodhf-published-sha256

What is recorded

Publisherqwen
OriginRecorded as a covered nation under 10 U.S.C. § 4872(f) — the PRC, Russia, Iran or North Korea. The registry did not record which.
LicenceApache-2.0
Parameters125,086,497,008 (125.1B)
Context window262K
Modalitytext+image+video->text
Native precisionBF16
Pulls recorded here360,013
Downloads reported upstream523,504

Reported scores

GPQA Diamond85.7
IFBench75.7
Humanity's Last Exam25.2

Reported by a third party and recorded here, not re-run by us. Source: artificialanalysis.ai/api/v2.

Measured on real hardware

None of these runs are ours. Each figure is the reporter’s, linked to their own evidence, under their own trust label. The number shown is single-stream decode — one user, concurrency 1 — because aggregate throughput across many users is a different quantity and reads as far faster than anyone will see.

Reported to local-ai-registry

HardwareEngineBuild measuredDecode, 1 streamLabelReported by
GeForce RTX 3090 ×4 devicesvllmsafetensors AWQ152 tok/s @ 2,048candidatelocalmaxxing
RTX PRO 6000 Blackwellllama.cppQ4_K_P94.6 tok/s @ 32,768candidatelocalmaxxing
RTX PRO 6000 Blackwellllama.cppQ4_K_M88.2 tok/s @ 32,768candidatelocalmaxxing
Apple M5 Max 128GBmlxMLX 4bit55.9 tok/s @ 8,256candidateexo-postgres
Apple M3 Ultra 96GB 80-core GPUmlxMLX 4bit51.1 tok/s @ 8,256candidateexo-postgres
Apple M3 Ultra 96GB 60-core GPUmlxMLX 4bit50.4 tok/s @ 8,256candidateexo-postgres
GeForce RTX 3090 ×8 devicesvllmsafetensors GPTQ Int449.1 tok/s @ 8,192candidatelocalmaxxing
GeForce RTX 3090 ×8 devicesvllmsafetensors FP87.4 tok/s @ 8,192candidatelocalmaxxing
Measured outside 2K–32K context — decode speed depends heavily on context, so these are not comparable with the rows above
GeForce RTX 3090 ×4 devicesllama.cppUD-Q4_K_XL51.3 tok/s @ 65,536candidatelocalmaxxing
NVIDIA DGX Spark GB10vllmINT433.9 tok/s @ 131,072candidatelocalmaxxing

From local-ai-registry (MIT), commit 124959d: 25 runs across 8 machines, 0 validated — model revision and runtime pinned, launch accepted — and 25 candidate, their label for useful evidence without a reproducible-launch promise. 3 more not listed.

Measured by local.ai

HardwareBuild measuredDecode @ 8KMemory @ 8KDecode mode
MacBook Pro M5 Max 128GB · 40-core GPU4bit54.9 tok/s72.6 GBnot stated
Mac Studio M3 Ultra 96GB · 60-core GPU4bit50.8 tok/s72.6 GBnot stated

Measured by local.ai, not by us: 2 runs. Each names its engine pinned by image digest and the exact serve command, in the full record.

Elsewhere in the registry

Registry record as of 2026-06-11. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.