architectures · deepseek-ai

DeepSeek-V4 (MLA + MoE)

Continues the DeepSeek line's pairing of multi-head latent attention with sparse experts: low-rank query projection (q_lora_rank 1536) and a compressed key/value path, 384 routed experts with 6 active per token plus one shared expert always on, across 61 layers. As with V3, MLA means the per-head KV-cache formula does not apply and the registry returns no estimate rather than a wrong one. Declared as model_type `deepseek_v4`.

What this registry has checked

Quarantine. We hold the publisher’s own published hash for these weights and have recorded it against this entry. We have not downloaded and hashed the weights ourselves, so the hash is the publisher’s claim, not our measurement.

What is recorded

Publisherdeepseek-ai
OriginChina
LicenceDeepSeek-License

Elsewhere in the registry

Registry record as of 2026-08-14. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.