architectures · allenai
Dense decoder-only transformer that interleaves sliding-window attention (4096-token window) with periodic full-attention layers, so most layers keep a bounded KV cache while the full-attention layers preserve long-range reach. Grouped-query attention; 64K context. Continues the OLMo line's defining property of releasing data and training code alongside weights. Declared as model_type `olmo3`.
| Publisher | allenai |
|---|---|
| Origin | United States |
| Licence | Apache-2.0 |
Registry record as of 2026-08-14. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.