protocols · ggml-org

GGUF

Single-file model container used by the llama.cpp family: tensors plus a key/value metadata header carrying architecture, tokenizer and quantization scheme, so a runtime can load a model without separate config or tokenizer files. Its per-tensor quantization types are what k-quant names like Q4_K_M refer to, and are why a real GGUF file's size differs from a naive params x bits calculation — embedding and output tensors are commonly kept at higher precision than the body.

What this registry has checked

Quarantine. We hold the publisher’s own published hash for these weights and have recorded it against this entry. We have not downloaded and hashed the weights ourselves, so the hash is the publisher’s claim, not our measurement.

What is recorded

Publisherggml-org
OriginInternational
LicenceMIT
Pulls recorded here2

Elsewhere in the registry

Registry record as of 2026-08-14. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.