models · qwen
Qwen3.6 dense 27.8B. 262K context, 150 tok/s. Arena 1138. GPQA Diamond 87.8%, LiveCodeBench 75.8%. Strong dense model at 28B scale.
| Weight file SHA-256 | 20e9d0e734a8fb4e4f9a2ddd9aead33317e737c2b751bc601b6bec0173ab5237 |
|---|---|
| Method | hf-published-sha256 |
| Publisher | qwen |
|---|---|
| Origin | Recorded as a covered nation under 10 U.S.C. § 4872(f) — the PRC, Russia, Iran or North Korea. The registry did not record which. |
| Licence | Apache-2.0 |
| Parameters | 27,781,427,952 (27.8B) |
| Context window | 262K |
| Modality | text+image+video->text |
| Native precision | BF16 |
| Pulls recorded here | 420,014 |
| Downloads reported upstream | 4,048,692 |
| GPQA Diamond | 84.2 |
|---|---|
| IFBench | 67.6 |
| Humanity's Last Exam | 23.1 |
Reported by a third party and recorded here, not re-run by us. Source: artificialanalysis.ai/api/v2.
| Hardware | Engine | Build measured | Decode, 1 stream | Label | Reported by |
|---|---|---|---|---|---|
| Radeon AI PRO R9700 | hipfire | MQ4 | 278 tok/s @ 2,048 | candidate | localmaxxing |
| Radeon RX 7900 XTX | hipfire | MQ4 | 252 tok/s @ 2,048 | candidate | localmaxxing |
| GeForce RTX 3090 Ti | llama.cpp | UD-Q4_K_XL | 156 tok/s @ 4,096 | candidate | localmaxxing |
| GeForce RTX 3090 | llama.cpp | UD-Q4_K_XL | 150 tok/s @ 2,048 | candidate | localmaxxing |
| RTX PRO 6000 Blackwell | llama.cpp | Q4_0 | 144 tok/s @ 4,445 | candidate | localmaxxing |
| RTX PRO 6000 Blackwell | vllm | NVFP4 | 140 tok/s @ 2,048 | candidate | localmaxxing |
| Ryzen AI Max+ 395 | hipfire | MQ4 | 112 tok/s @ 2,048 | candidate | localmaxxing |
| GeForce RTX 5090 ×2 devices | vllm | safetensors NVFP4 | 85.3 tok/s @ 8,192 | candidate | localmaxxing |
| GeForce RTX 5090 | llama.cpp | safetensors IQ4_XS | 79.2 tok/s @ 8,192 | candidate | localmaxxing |
| GeForce RTX 3090 ×2 devices | vllm | safetensors FP8 | 77.8 tok/s @ 8,192 | candidate | localmaxxing |
From local-ai-registry (MIT), commit 124959d: 306 runs across 32 machines, 10 validated — model revision and runtime pinned, launch accepted — and 296 candidate, their label for useful evidence without a reproducible-launch promise. 198 more not listed. 3 report figures that contradict each other — prefill below decode, which parallel prefill essentially never is — and are left to the source rather than ranked here.
| Hardware | Build measured | Decode @ 8K | Memory @ 8K | Decode mode |
|---|---|---|---|---|
| NVIDIA RTX PRO 6000 Blackwell 96GB | Q4_K_M | 74.4 tok/s | 19.5 GB | ordinary |
| NVIDIA GeForce RTX 5090 32GB | NVFP4 | 73.9 tok/s | 28.2 GB | ordinary |
| NVIDIA RTX PRO 6000 Blackwell 96GB | NVFP4 | 73.0 tok/s | 90.2 GB | ordinary |
| NVIDIA GeForce RTX 5090 32GB | NVFP4 | 65.8 tok/s | 28.3 GB | ordinary |
| NVIDIA RTX PRO 6000 Blackwell 96GB | NVFP4 | 64.7 tok/s | 92.0 GB | ordinary |
| NVIDIA RTX 6000 Ada 48GB | Q4_K_M | 44.5 tok/s | 22.7 GB | ordinary |
| NVIDIA RTX 6000 Ada 48GB | NVFP4 | 38.0 tok/s | 43.4 GB | ordinary |
| Mac Studio M3 Ultra 96GB · 60-core GPU | 4bit | 28.9 tok/s | 23.2 GB | ordinary |
Measured by local.ai, not by us: 112 runs, 104 more not listed. Each names its engine pinned by image digest and the exact serve command, in the full record.
Registry record as of 2026-06-11. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.