models · meta-llama

Llama-3.2-3B-Instruct

Compact 3B instruct model for on-device and edge deployment with 128K context. Distilled from the larger Llama 3.1 models.

Compare Verify

What this registry has checked

Quarantine. We hold the publisher’s own published hash for these weights and have recorded it against this entry. We have not downloaded and hashed the weights ourselves, so the hash is the publisher’s claim, not our measurement.
Weight file SHA-2569b1ab57fbb64db069c8d9839fbfb1949c89281dd92703605f8e765bbc317c260
Methodself-download-failed

What is recorded

Publishermeta-llama
OriginUnited States
LicenceLlama-3.2
Publisher’s own licence statementllama3.2
Parameters3,212,749,824 (3.21B)
Context window128K
Modalitytext->text
Native precisionBF16
Pulls recorded here224,014
Downloads reported upstream1,690,943

Measured on real hardware

None of these runs are ours. Each figure is the reporter’s, linked to their own evidence, under their own trust label. The number shown is single-stream decode — one user, concurrency 1 — because aggregate throughput across many users is a different quantity and reads as far faster than anyone will see.

Reported to local-ai-registry

HardwareEngineBuild measuredDecode, 1 streamLabelReported by
GeForce RTX 3060llama.cppQ4_K_M128 tok/s @ 512candidatelocalmaxxing
Apple M1 Max 64GBllama.cppBF1650.9 tok/s @ 640candidatelocalmaxxing

From local-ai-registry (MIT), commit 124959d: 2 runs across 2 machines, 0 validated — model revision and runtime pinned, launch accepted — and 2 candidate, their label for useful evidence without a reproducible-launch promise.

Elsewhere in the registry

Registry record as of 2026-06-12. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.