architectures · tiiuae
Decoder-only transformer using multi-query attention — one shared key/value head across all query heads — which shrinks the KV cache substantially at long context and was the design's headline inference property. Parallel attention and feed-forward blocks; rotary position embeddings. Trained on the RefinedWeb corpus, published alongside the models. Declared as model_type `falcon`.
| Publisher | tiiuae |
|---|---|
| Origin | United Arab Emirates |
| Licence | Apache-2.0 |
Registry record as of 2026-08-14. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.