architectures · mistralai
Dense decoder-only transformer combining grouped-query attention with sliding-window attention, so each layer attends over a fixed window rather than the full sequence and the KV cache is bounded by the window instead of the context length. Distinct from the Mixtral sparse-MoE line, which shares the backbone but routes through expert feed-forward blocks. Declared as model_type `mistral` in config.json.
| Publisher | mistralai |
|---|---|
| Origin | France |
| Licence | Apache-2.0 |
Registry record as of 2026-08-14. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.