models · qwen
Qwen3 compact MoE. 30B total / 3B active. Low latency (285ms TTFT). Hybrid thinking mode for efficient reasoning on-device.
| Weight file SHA-256 | 505256bf25d85362d7933843ab7f888b0dba243e2bc71728ac1f88ac0d724edc |
|---|---|
| Method | hf-published-sha256 |
| Publisher | qwen |
|---|---|
| Origin | Recorded as a covered nation under 10 U.S.C. § 4872(f) — the PRC, Russia, Iran or North Korea. The registry did not record which. |
| Licence | Apache-2.0 |
| Parameters | 30,532,122,624 (30.5B) |
| Context window | 128K |
| Modality | text->text |
| Architecture | qwen3_moe |
| Layers | 48 |
| Hidden size | 2048 |
| Attention | GQA (32 heads, 4 KV) |
| Native precision | BF16 |
| Pulls recorded here | 890,032 |
| Downloads reported upstream | 1,859,450 |
| MMLU-Pro | 71 |
|---|---|
| GPQA Diamond | 51.5 |
| AIME 2025 | 21.7 |
| MATH-500 | 86.3 |
| LiveCodeBench | 32.2 |
| IFBench | 31.9 |
| Humanity's Last Exam | 4.6 |
Reported by a third party and recorded here, not re-run by us. Source: artificialanalysis.ai/api/v2.
| Hardware | Engine | Build measured | Decode, 1 stream | Label | Reported by |
|---|---|---|---|---|---|
| GeForce RTX 3060 | llama.cpp | Q4_K_M | 116 tok/s @ 2,048 | candidate | localmaxxing |
| Apple M3 Ultra 512GB | llama.cpp | Q4_K_M | 102 tok/s @ 8,192 | candidate | localmaxxing |
| GeForce RTX 3060 | ollama | Q4_K_M | 90.0 tok/s @ 2,048 | candidate | localmaxxing |
| Apple M1 Max 64GB | ollama | Q4_K_M | 62.6 tok/s @ 2,048 | candidate | localmaxxing |
| GeForce RTX 3080 | ollama | Q4_K_M | 31.0 tok/s @ 2,048 | candidate | localmaxxing |
From local-ai-registry (MIT), commit 124959d: 5 runs across 4 machines, 0 validated — model revision and runtime pinned, launch accepted — and 5 candidate, their label for useful evidence without a reproducible-launch promise.
Registry record as of 2026-06-11. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.