architectures · mistralai

Mistral (Dense)

Dense decoder-only transformer combining grouped-query attention with sliding-window attention, so each layer attends over a fixed window rather than the full sequence and the KV cache is bounded by the window instead of the context length. Distinct from the Mixtral sparse-MoE line, which shares the backbone but routes through expert feed-forward blocks. Declared as model_type `mistral` in config.json.

What this registry has checked

Quarantine. We hold the publisher’s own published hash for these weights and have recorded it against this entry. We have not downloaded and hashed the weights ourselves, so the hash is the publisher’s claim, not our measurement.

What is recorded

Publishermistralai
OriginFrance
LicenceApache-2.0

Elsewhere in the registry

Registry record as of 2026-08-14. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.