models · deepseek-ai
DeepSeek V3.2 non-thinking mode. 685B MoE, 131K context, 318 tok/s. Arena 980. High-throughput dense-equivalent inference without chain-of-thought overhead.
| Weight file SHA-256 | fd2fe54fb74c820540b4fc21f59b49123d5f6030a24cebbeffd36c92ad253838 |
|---|---|
| Method | hf-published-sha256 |
| Publisher | deepseek-ai |
|---|---|
| Origin | Recorded as a covered nation under 10 U.S.C. § 4872(f) — the PRC, Russia, Iran or North Korea. The registry did not record which. |
| Licence | MIT |
| Parameters | 685,396,921,376 (685.4B) |
| Context window | 131K |
| Modality | text->text |
| Architecture | deepseek_v3 |
| Layers | 61 |
| Hidden size | 7168 |
| Attention | MLA (128 heads, 128 KV) |
| Native precision | F8_E4M3 |
| Pulls recorded here | 870,035 |
| Downloads reported upstream | 1,959,027 |
Registry record as of 2026-09-15. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.