protocols · ggml-org
Single-file model container used by the llama.cpp family: tensors plus a key/value metadata header carrying architecture, tokenizer and quantization scheme, so a runtime can load a model without separate config or tokenizer files. Its per-tensor quantization types are what k-quant names like Q4_K_M refer to, and are why a real GGUF file's size differs from a naive params x bits calculation — embedding and output tensors are commonly kept at higher precision than the body.
| Publisher | ggml-org |
|---|---|
| Origin | International |
| Licence | MIT |
| Pulls recorded here | 2 |
Registry record as of 2026-08-14. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.