models · Qwen
Catalogued from llm-stats Open LLM Leaderboard at Quarantine — unverified, pending attestation.
| Method | hf-published-sha256 |
|---|
| Publisher | Qwen |
|---|---|
| Origin | Recorded as a covered nation under 10 U.S.C. § 4872(f) — the PRC, Russia, Iran or North Korea. The registry did not record which. |
| Licence | Apache-2.0 |
| Parameters | 9,653,104,368 (9.65B) |
| Context window | 262,144 tokens |
| Modality | text+image+video->text |
| Native precision | BF16 |
| Pulls recorded here | 6 |
| Downloads reported upstream | 9,468,124 |
| GPQA Diamond | 80.6 |
|---|---|
| IFBench | 66.7 |
| Humanity's Last Exam | 14.9 |
Reported by a third party and recorded here, not re-run by us. Source: artificialanalysis.ai/api/v2.
| Hardware | Engine | Build measured | Decode, 1 stream | Label | Reported by |
|---|---|---|---|---|---|
| GeForce RTX 5090 | llama.cpp | Q6_K | 170 tok/s @ 2,048 | candidate | localmaxxing |
| GeForce RTX 5070 Ti | llama.cpp | Q4_K_S | 124 tok/s @ 8,192 | candidate | localmaxxing |
| GeForce RTX 3090 Ti | llama.cpp | UD-Q4_K_XL | 119 tok/s @ 8,192 | candidate | localmaxxing |
| GeForce RTX 3090 ×2 devices | llama.cpp | Q4_K_S | 110 tok/s @ 8,192 | candidate | localmaxxing |
| H200 141GB | ollama | GGUF Q4_K_M | 103 tok/s @ 32,768 | candidate | 0xsero |
| Radeon RX 9070 XT | llama.cpp | Q4_K_M | 87.5 tok/s @ 4,096 | candidate | localmaxxing |
| Apple M3 Ultra 96GB 80-core GPU | llama.cpp | GGUF Q4_K_M | 81.6 tok/s @ 8,256 | candidate | exo-postgres |
| GeForce RTX 5080 | llama.cpp | Q8_0 | 79.6 tok/s @ 4,096 | candidate | localmaxxing |
| Apple M5 Max 128GB | llama.cpp | Q4_K_M | 79.3 tok/s @ 2,048 | candidate | localmaxxing |
| Apple M5 Max 128GB | llama.cpp | GGUF Q4_K_M | 78.6 tok/s @ 8,256 | candidate | exo-postgres |
From local-ai-registry (MIT), commit 124959d: 99 runs across 51 machines, 19 validated — model revision and runtime pinned, launch accepted — and 80 candidate, their label for useful evidence without a reproducible-launch promise. 48 more not listed. 21 report figures that contradict each other — prefill below decode, which parallel prefill essentially never is — and are left to the source rather than ranked here.
| Hardware | Build measured | Decode @ 8K | Memory @ 8K | Decode mode |
|---|---|---|---|---|
| NVIDIA GeForce RTX 5090 32GB | Q4_K_M | 212 tok/s | 10.3 GB | ordinary |
| NVIDIA RTX PRO 6000 Blackwell 96GB | Q4_K_M | 204 tok/s | 6.9 GB | ordinary |
| NVIDIA RTX PRO 6000 Blackwell 96GB | FP8 | 140 tok/s | 90.0 GB | ordinary |
| NVIDIA GeForce RTX 5090 32GB | FP8 | 139 tok/s | 28.4 GB | ordinary |
| NVIDIA GeForce RTX 4090 24GB | Q4_K_M | 129 tok/s | 8.0 GB | ordinary |
| NVIDIA RTX 6000 Ada 48GB | Q4_K_M | 127 tok/s | 14.6 GB | ordinary |
| Mac Studio M3 Ultra 96GB · 80-core GPU | 4bit | 94.0 tok/s | 7.5 GB | ordinary |
| Mac Studio M3 Ultra 96GB · 60-core GPU | 4bit | 90.9 tok/s | 7.5 GB | ordinary |
Measured by local.ai, not by us: 133 runs, 125 more not listed. Each names its engine pinned by image digest and the exact serve command, in the full record.
Registry record as of 2026-09-26. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.