architectures · allenai

OLMo 3

Dense decoder-only transformer that interleaves sliding-window attention (4096-token window) with periodic full-attention layers, so most layers keep a bounded KV cache while the full-attention layers preserve long-range reach. Grouped-query attention; 64K context. Continues the OLMo line's defining property of releasing data and training code alongside weights. Declared as model_type `olmo3`.

What this registry has checked

Quarantine. We hold the publisher’s own published hash for these weights and have recorded it against this entry. We have not downloaded and hashed the weights ourselves, so the hash is the publisher’s claim, not our measurement.

What is recorded

Publisherallenai
OriginUnited States
LicenceApache-2.0

Elsewhere in the registry

Registry record as of 2026-08-14. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.