models · google
Gemma 3 12B instruction-tuned with multimodal input support.
| Weight file SHA-256 | 094f35c2a78344d1ff4e6d16016a1b37ebb377173d68bcfa0ce2b103544d7b46 |
|---|---|
| Method | hf-published-sha256 |
| Publisher | |
|---|---|
| Origin | United States |
| Licence | Gemma-TOS |
| Publisher’s own licence statement | gemma |
| Parameters | 12,187,325,040 (12.2B) |
| Context window | 131,072 tokens |
| Modality | text+image->text |
| Native precision | BF16 |
| Pulls recorded here | 48,019 |
| Downloads reported upstream | 681,214 |
| GPQA Diamond | 56.2 |
|---|
Reported by a third party and recorded here, not re-run by us.
| Hardware | Engine | Build measured | Decode, 1 stream | Label | Reported by |
|---|---|---|---|---|---|
| GeForce RTX 3080 | ollama | Q4_K_M | 76.0 tok/s @ 2,048 | candidate | localmaxxing |
| GeForce RTX 3060 | ollama | Q4_K_M | 40.0 tok/s @ 2,048 | candidate | localmaxxing |
| Measured outside 2K–32K context — decode speed depends heavily on context, so these are not comparable with the rows above | |||||
| GeForce RTX 3090 | llama.cpp | Q4_K_M | 66.0 tok/s @ 512 | candidate | localmaxxing |
| GeForce RTX 3060 ×2 devices | llama.cpp | Q4_K_M | 39.4 tok/s @ 512 | candidate | localmaxxing |
From local-ai-registry (MIT), commit 124959d: 4 runs across 3 machines, 0 validated — model revision and runtime pinned, launch accepted — and 4 candidate, their label for useful evidence without a reproducible-launch promise.
Registry record as of 2026-06-11. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.