architectures · deepseek-ai
Continues the DeepSeek line's pairing of multi-head latent attention with sparse experts: low-rank query projection (q_lora_rank 1536) and a compressed key/value path, 384 routed experts with 6 active per token plus one shared expert always on, across 61 layers. As with V3, MLA means the per-head KV-cache formula does not apply and the registry returns no estimate rather than a wrong one. Declared as model_type `deepseek_v4`.
| Publisher | deepseek-ai |
|---|---|
| Origin | China |
| Licence | DeepSeek-License |
Registry record as of 2026-08-14. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.