models · qwen
Qwen3 32B dense model with hybrid thinking/non-thinking mode. Strong coding and math.
| Weight file SHA-256 | dfbadf618f9612482bbed4bee379953b7094fc979770e5b1f8d55038f7457b1a |
|---|---|
| Method | hf-published-sha256 |
| Publisher | qwen |
|---|---|
| Origin | China |
| Licence | Apache-2.0 |
| Parameters | 32,762,123,264 (32.8B) |
| Context window | 131,072 tokens |
| Modality | text->text |
| Architecture | qwen3 |
| Layers | 64 |
| Hidden size | 5120 |
| Attention | GQA (64 heads, 8 KV) |
| Native precision | BF16 |
| Pulls recorded here | 94,021 |
| Downloads reported upstream | 4,674,951 |
| MMLU-Pro | 72.7 |
|---|---|
| GPQA Diamond | 53.5 |
| AIME 2025 | 19.7 |
| MATH-500 | 86.9 |
| LiveCodeBench | 28.8 |
| IFBench | 31.5 |
| Humanity's Last Exam | 4.1 |
Reported by a third party and recorded here, not re-run by us. Source: artificialanalysis.ai/api/v2.
| Hardware | Engine | Build measured | Decode, 1 stream | Label | Reported by |
|---|---|---|---|---|---|
| Apple M1 Max 64GB | ollama | Q4_K_M | 9.0 tok/s @ 2,048 | candidate | localmaxxing |
From local-ai-registry (MIT), commit 124959d: 2 runs across 2 machines, 0 validated — model revision and runtime pinned, launch accepted — and 2 candidate, their label for useful evidence without a reproducible-launch promise.
| Hardware | Build measured | Decode @ 8K | Memory @ 8K | Decode mode |
|---|---|---|---|---|
| NVIDIA RTX PRO 6000 Blackwell 96GB | — | 22.4 tok/s | 89.9 GB | not stated |
| NVIDIA DGX Spark 128GB | — | 3.5 tok/s | 113.4 GB | not stated |
Measured by local.ai, not by us: 2 runs. Each names its engine pinned by image digest and the exact serve command, in the full record.
Registry record as of 2026-06-11. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.