models · deepseek-ai
text-generation model harvested from Hugging Face (trending). Catalogued at Quarantine — unverified, pending attestation.
| Method | hf-published-sha256 |
|---|
| Publisher | deepseek-ai |
|---|---|
| Origin | Recorded as a covered nation under 10 U.S.C. § 4872(f) — the PRC, Russia, Iran or North Korea. The registry did not record which. |
| Licence | mit |
| Parameters | 304,180,418,494 (304.2B) |
| Context window | 1,310,720 tokens |
| Modality | text->text |
| Architecture | deepseek_v4 |
| Layers | 43 |
| Hidden size | 4096 |
| Attention | GQA (64 heads, 1 KV) |
| Native precision | I8 |
| Pulls recorded here | 15,372 |
| Downloads reported upstream | 4,469,509 |
| GPQA Diamond | 90.8 |
|---|---|
| Humanity's Last Exam | 38.6 |
Reported by a third party and recorded here, not re-run by us. Source: artificialanalysis.ai/api/v2.
| Hardware | Engine | Build measured | Decode, 1 stream | Label | Reported by |
|---|---|---|---|---|---|
| Radeon AI PRO R9700 | llama.cpp | MXFP4 | 52.0 tok/s @ 32,768 | candidate | localmaxxing |
| Radeon AI PRO R9700 ×3 devices | llama.cpp | AD-IQ1_M | 48.5 tok/s @ 32,768 | candidate | localmaxxing |
| NVIDIA DGX Spark GB10 ×2 devices | vllm | ModelOpt NVFP4 | 43.7 tok/s @ 8,192 | candidate | mialabs |
| GeForce RTX 3090 | llama.cpp | IQ2_XSS | 20.7 tok/s @ 2,048 | candidate | localmaxxing |
| GeForce RTX 5090 | llama.cpp | UD-IQ1_S | 18.0 tok/s @ 2,048 | candidate | localmaxxing |
| GeForce RTX 5090 | llama.cpp | IQ1_S | 17.5 tok/s @ 2,048 | candidate | localmaxxing |
| GeForce RTX 5090 | llama.cpp | Q8_K_XL | 10.4 tok/s @ 2,048 | candidate | localmaxxing |
| Measured outside 2K–32K context — decode speed depends heavily on context, so these are not comparable with the rows above | |||||
| RTX PRO 6000 Blackwell ×2 devices | vllm | safetensors FP8_E4M3 | 276 tok/s @ 262,144 | candidate | localmaxxing |
| RTX PRO 6000 Blackwell ×2 devices | sglang | safetensors FP8 | 57.8 tok/s @ 50 | candidate | ppickle1989 |
| GeForce RTX 3090 ×7 devices | llama.cpp | UD-IQ4_XS | 50.8 tok/s @ 262,144 | candidate | localmaxxing |
From local-ai-registry (MIT), commit 124959d: 73 runs across 6 machines, 0 validated — model revision and runtime pinned, launch accepted — and 73 candidate, their label for useful evidence without a reproducible-launch promise. 5 more not listed.
| Hardware | Build measured | Decode @ 8K | Memory @ 8K | Decode mode |
|---|---|---|---|---|
| MacBook Pro M5 Max 128GB · 40-core GPU | UD-Q2_K_XL | 29.9 tok/s | 99.1 GB | ordinary |
| NVIDIA DGX Spark 128GB | UD-IQ2_M | 19.9 tok/s | 93.1 GB | ordinary |
| NVIDIA DGX Spark 128GB | UD-IQ3_XXS | 18.8 tok/s | 105.5 GB | ordinary |
| NVIDIA DGX Spark 128GB | UD-Q2_K_XL | 17.5 tok/s | 98.8 GB | ordinary |
| Mac Studio M3 Ultra 96GB · 80-core GPU | UD-Q2_K_XL | 7.1 tok/s | 99.1 GB | ordinary |
Measured by local.ai, not by us: 5 runs. Each names its engine pinned by image digest and the exact serve command, in the full record.
Registry record as of 2026-09-01. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.