architectures · sarvamai
Multi-head latent attention combined with sparse experts: keys and values are projected into a shared 512-dimension latent (kv_lora_rank 512) rather than cached per head, alongside 128 experts with 8 routed per token. Because MLA compresses KV into a shared latent, the standard per-head KV-cache formula does not apply to this architecture and the registry declines to estimate rather than reporting a wrong number. Declared as model_type `sarvam_mla`.
| Publisher | sarvamai |
|---|---|
| Origin | India |
| Licence | Apache-2.0 |
Registry record as of 2026-08-14. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.