architectures · openbmb
Dense decoder-only transformer built for very long context: 512K positions with an aggressive 16:1 grouped-query ratio (32 query heads against 2 key/value heads), which is what keeps the KV cache tractable at that length — cache size scales with key/value heads, not query heads. Declared as model_type `minicpm_sala`.
| Publisher | openbmb |
|---|---|
| Origin | China |
| Licence | Apache-2.0 |
Registry record as of 2026-08-14. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.