architectures · inclusionai
Sparse mixture-of-experts backbone behind the Ling 3.0 family, with 128 experts and 8 routed per token, and a low-rank compressed key/value projection (kv_lora_rank 512) rather than a full per-head KV cache. Declared as model_type `bailing_hybrid`, architecture class BailingMoeV3.
| Publisher | inclusionai |
|---|---|
| Origin | China |
| Licence | Apache-2.0 |
Registry record as of 2026-08-14. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.