models · qwen
Qwen3 14B dense instruct model. Efficient mid-range option with strong reasoning and multilingual performance.
| Weight file SHA-256 | 245877f20829ecc834373c6fc1e86dfb6bc6a699684dbb2967033d4f7a4fc786 |
|---|---|
| Method | hf-published-sha256 |
| Publisher | qwen |
|---|---|
| Origin | Recorded as a covered nation under 10 U.S.C. § 4872(f) — the PRC, Russia, Iran or North Korea. The registry did not record which. |
| Licence | Apache-2.0 |
| Parameters | 14,768,307,200 (14.8B) |
| Context window | 128K |
| Modality | text->text |
| Architecture | qwen3 |
| Layers | 40 |
| Hidden size | 5120 |
| Attention | GQA (40 heads, 8 KV) |
| Native precision | BF16 |
| Pulls recorded here | 750,019 |
| Downloads reported upstream | 1,584,558 |
| MMLU-Pro | 67.5 |
|---|---|
| GPQA Diamond | 47 |
| AIME 2025 | 58 |
| MATH-500 | 87.1 |
| LiveCodeBench | 28 |
| IFBench | 23.9 |
| Humanity's Last Exam | 4.1 |
Reported by a third party and recorded here, not re-run by us. Source: artificialanalysis.ai/api/v2.
| Hardware | Engine | Build measured | Decode, 1 stream | Label | Reported by |
|---|---|---|---|---|---|
| RTX PRO 6000 Blackwell | vllm | safetensors FP8 | 90.8 tok/s @ 32,768 | candidate | localmaxxing |
| GeForce RTX 3080 | llama.cpp | Q4_K_M | 75.0 tok/s @ 2,048 | candidate | localmaxxing |
| GeForce RTX 5070 | llama.cpp | Q4_K_M | 66.6 tok/s @ 2,048 | candidate | localmaxxing |
| GeForce RTX 3060 | llama.cpp | Q4_K_M | 37.0 tok/s @ 2,048 | candidate | localmaxxing |
| GeForce RTX 3060 | ollama | Q4_K_M | 35.0 tok/s @ 2,048 | candidate | localmaxxing |
| Apple M1 Max 64GB | ollama | Q4_K_M | 24.8 tok/s @ 2,048 | candidate | localmaxxing |
| Measured outside 2K–32K context — decode speed depends heavily on context, so these are not comparable with the rows above | |||||
| GeForce RTX 3090 | llama.cpp | Q4_K_M | 51.7 tok/s @ 1,024 | candidate | localmaxxing |
| GeForce RTX 3090 | llama.cpp | Q4_K_M | 51.7 tok/s @ 512 | candidate | localmaxxing |
| GeForce RTX 3060 ×2 devices | llama.cpp | Q4_K_M | 35.8 tok/s @ 512 | candidate | localmaxxing |
| GeForce RTX 3060 | llama.cpp | Q4_K_M | 34.6 tok/s @ 512 | candidate | localmaxxing |
From local-ai-registry (MIT), commit 124959d: 10 runs across 6 machines, 0 validated — model revision and runtime pinned, launch accepted — and 10 candidate, their label for useful evidence without a reproducible-launch promise.
Registry record as of 2026-06-11. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.