architectures · sarvamai
Sparse mixture-of-experts transformer with 128 experts and 6 routed per token, using grouped-query attention over 64 query heads against 4 key/value heads. Distinct from the MLA line, which compresses KV into a shared latent instead. Declared as model_type `sarvam_moe`.
| Publisher | sarvamai |
|---|---|
| Origin | India |
| Licence | Apache-2.0 |
Registry record as of 2026-08-14. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.