models · deepseek-ai
DeepSeek-V4 efficient inference variant. Optimized for low-cost, high-throughput deployment. 284B parameters, 1M context window.
| Weight file SHA-256 | 94a4467bac50616db8e2fdadc2bb1d49aa98c11b97d51e492d4d83badf82df36 |
|---|---|
| Method | hf-published-sha256 |
| Publisher | deepseek-ai |
|---|---|
| Origin | Recorded as a covered nation under 10 U.S.C. § 4872(f) — the PRC, Russia, Iran or North Korea. The registry did not record which. |
| Licence | MIT |
| Parameters | 290,944,616,402 (290.9B) |
| Context window | 1M |
| Modality | text->text |
| Architecture | deepseek_v4 |
| Layers | 43 |
| Hidden size | 4096 |
| Attention | GQA (64 heads, 1 KV) |
| Native precision | I8 |
| Pulls recorded here | 670,019 |
| Downloads reported upstream | 1,750,419 |
| GPQA Diamond | 89.4 |
|---|---|
| IFBench | 79.2 |
| Humanity's Last Exam | 34.8 |
Reported by a third party and recorded here, not re-run by us. Source: artificialanalysis.ai/api/v2.
| Hardware | Engine | Build measured | Decode, 1 stream | Label | Reported by |
|---|---|---|---|---|---|
| Radeon AI PRO R9700 ×4 devices | hipfire | MQ2R | 54.3 tok/s @ 2,052 | candidate | localmaxxing |
| RTX PRO 6000 Blackwell ×4 devices | sglang | safetensors FP8 | 37.6 tok/s @ 8,192 | candidate | 0xsero |
| Apple M5 Max 128GB | llama.cpp | GGUF UD_Q2_K_XL | 32.2 tok/s @ 8,256 | candidate | exo-postgres |
| Apple M5 Max 128GB | llama.cpp | IQ2_XXS (asym MoE) | 26.3 tok/s @ 32,768 | candidate | localmaxxing |
| Apple M5 Max 128GB | llama.cpp | GGUF IQ2XXS | 25.5 tok/s @ 8,256 | candidate | exo-postgres |
| Apple M3 Ultra 96GB 80-core GPU | llama.cpp | GGUF IQ2XXS | 22.5 tok/s @ 8,256 | candidate | exo-postgres |
| Apple M4 Max 128GB | llama.cpp | GGUF IQ2XXS | 20.1 tok/s @ 8,256 | candidate | exo-postgres |
| Apple M4 Max 128GB | llama.cpp | GGUF IQ2XXS | 20.1 tok/s @ 8,256 | candidate | exo-postgres |
| Apple M5 Max 128GB | llama.cpp | GGUF UD-Q2_K_XL | 13.0 tok/s @ 8,256 | candidate | exo-postgres |
| Apple M5 Max 128GB | llama.cpp | GGUF IQ2XXS | 12.7 tok/s @ 8,256 | candidate | exo-postgres |
From local-ai-registry (MIT), commit 124959d: 82 runs across 13 machines, 0 validated — model revision and runtime pinned, launch accepted — and 82 candidate, their label for useful evidence without a reproducible-launch promise. 16 more not listed.
| Hardware | Build measured | Decode @ 8K | Memory @ 8K | Decode mode |
|---|---|---|---|---|
| MacBook Pro M5 Max 128GB · 40-core GPU | UD-Q2_K_XL | 29.9 tok/s | 99.1 GB | ordinary |
| NVIDIA DGX Spark 128GB | UD-Q2_K_XL | 17.5 tok/s | 98.8 GB | ordinary |
| Mac Studio M3 Ultra 96GB · 80-core GPU | UD-Q2_K_XL | 7.1 tok/s | 99.1 GB | ordinary |
Measured by local.ai, not by us: 3 runs. Each names its engine pinned by image digest and the exact serve command, in the full record.
Registry record as of 2026-08-13. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.