models · qwen
Qwen3.5 MoE 122B total / 10B active. 262K context, 129 tok/s. Arena 750. GPQA Diamond 86.6%. Strong MoE efficiency at near-70B inference cost.
| Weight file SHA-256 | 9c064915c4742a03e8d5e317e3841f521e9bf51a5cedfe3829d4e173b13555a5 |
|---|---|
| Method | hf-published-sha256 |
| Publisher | qwen |
|---|---|
| Origin | Recorded as a covered nation under 10 U.S.C. § 4872(f) — the PRC, Russia, Iran or North Korea. The registry did not record which. |
| Licence | Apache-2.0 |
| Parameters | 125,086,497,008 (125.1B) |
| Context window | 262K |
| Modality | text+image+video->text |
| Native precision | BF16 |
| Pulls recorded here | 360,013 |
| Downloads reported upstream | 523,504 |
| GPQA Diamond | 85.7 |
|---|---|
| IFBench | 75.7 |
| Humanity's Last Exam | 25.2 |
Reported by a third party and recorded here, not re-run by us. Source: artificialanalysis.ai/api/v2.
| Hardware | Engine | Build measured | Decode, 1 stream | Label | Reported by |
|---|---|---|---|---|---|
| GeForce RTX 3090 ×4 devices | vllm | safetensors AWQ | 152 tok/s @ 2,048 | candidate | localmaxxing |
| RTX PRO 6000 Blackwell | llama.cpp | Q4_K_P | 94.6 tok/s @ 32,768 | candidate | localmaxxing |
| RTX PRO 6000 Blackwell | llama.cpp | Q4_K_M | 88.2 tok/s @ 32,768 | candidate | localmaxxing |
| Apple M5 Max 128GB | mlx | MLX 4bit | 55.9 tok/s @ 8,256 | candidate | exo-postgres |
| Apple M3 Ultra 96GB 80-core GPU | mlx | MLX 4bit | 51.1 tok/s @ 8,256 | candidate | exo-postgres |
| Apple M3 Ultra 96GB 60-core GPU | mlx | MLX 4bit | 50.4 tok/s @ 8,256 | candidate | exo-postgres |
| GeForce RTX 3090 ×8 devices | vllm | safetensors GPTQ Int4 | 49.1 tok/s @ 8,192 | candidate | localmaxxing |
| GeForce RTX 3090 ×8 devices | vllm | safetensors FP8 | 7.4 tok/s @ 8,192 | candidate | localmaxxing |
| Measured outside 2K–32K context — decode speed depends heavily on context, so these are not comparable with the rows above | |||||
| GeForce RTX 3090 ×4 devices | llama.cpp | UD-Q4_K_XL | 51.3 tok/s @ 65,536 | candidate | localmaxxing |
| NVIDIA DGX Spark GB10 | vllm | INT4 | 33.9 tok/s @ 131,072 | candidate | localmaxxing |
From local-ai-registry (MIT), commit 124959d: 25 runs across 8 machines, 0 validated — model revision and runtime pinned, launch accepted — and 25 candidate, their label for useful evidence without a reproducible-launch promise. 3 more not listed.
| Hardware | Build measured | Decode @ 8K | Memory @ 8K | Decode mode |
|---|---|---|---|---|
| MacBook Pro M5 Max 128GB · 40-core GPU | 4bit | 54.9 tok/s | 72.6 GB | not stated |
| Mac Studio M3 Ultra 96GB · 60-core GPU | 4bit | 50.8 tok/s | 72.6 GB | not stated |
Measured by local.ai, not by us: 2 runs. Each names its engine pinned by image digest and the exact serve command, in the full record.
Registry record as of 2026-06-11. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.