models · meta-llama
Compact 3B instruct model for on-device and edge deployment with 128K context. Distilled from the larger Llama 3.1 models.
| Weight file SHA-256 | 9b1ab57fbb64db069c8d9839fbfb1949c89281dd92703605f8e765bbc317c260 |
|---|---|
| Method | self-download-failed |
| Publisher | meta-llama |
|---|---|
| Origin | United States |
| Licence | Llama-3.2 |
| Publisher’s own licence statement | llama3.2 |
| Parameters | 3,212,749,824 (3.21B) |
| Context window | 128K |
| Modality | text->text |
| Native precision | BF16 |
| Pulls recorded here | 224,014 |
| Downloads reported upstream | 1,690,943 |
| Hardware | Engine | Build measured | Decode, 1 stream | Label | Reported by |
|---|---|---|---|---|---|
| GeForce RTX 3060 | llama.cpp | Q4_K_M | 128 tok/s @ 512 | candidate | localmaxxing |
| Apple M1 Max 64GB | llama.cpp | BF16 | 50.9 tok/s @ 640 | candidate | localmaxxing |
From local-ai-registry (MIT), commit 124959d: 2 runs across 2 machines, 0 validated — model revision and runtime pinned, launch accepted — and 2 candidate, their label for useful evidence without a reproducible-launch promise.
Registry record as of 2026-06-12. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.