architectures · arcee-ai
Sparse mixture-of-experts transformer with a wide expert pool and narrow routing — 256 experts with 4 routed per token — combined with grouped-query attention and a 4096-token sliding window. The wide-pool/narrow-route ratio is the design's distinguishing choice: capacity scales with the pool while per-token compute stays near a much smaller dense model. Declared as model_type `afmoe`.
| Publisher | arcee-ai |
|---|---|
| Origin | United States |
| Licence | Apache-2.0 |
Registry record as of 2026-08-14. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.