models · google
Google Gemma 4 MoE variant. 26B total / 4B active parameters. 262K context. Efficient inference with 115 tok/s throughput.
| Weight file SHA-256 | 25f5f6945c06b525d613f31a32c53d449de0e004f03d75ab5ddeb7a4d5823a48 |
|---|---|
| Method | hf-published-sha256 |
| Publisher | |
|---|---|
| Origin | United States |
| Licence | Apache-2.0 |
| Parameters | 25,805,936,206 (26B-A4B) |
| Context window | 262K |
| Modality | text+image+video->text |
| Native precision | BF16 |
| Pulls recorded here | 180,015 |
| Downloads reported upstream | 48,151 |
| GPQA Diamond | 79.2 |
|---|---|
| IFBench | 72.4 |
| Humanity's Last Exam | 19.3 |
Reported by a third party and recorded here, not re-run by us. Source: artificialanalysis.ai/api/v2.
| Hardware | Engine | Build measured | Decode, 1 stream | Label | Reported by |
|---|---|---|---|---|---|
| GeForce RTX 5090 | llama.cpp | Q3_K_M | 256 tok/s @ 2,048 | candidate | localmaxxing |
| Intel Arc Pro B70 | llama.cpp | Q8_K_XL | 125 tok/s @ 32,768 | candidate | localmaxxing |
| Radeon PRO W7800 | llama.cpp | Q4_K_S | 117 tok/s @ 8,192 | validated | 0xsero |
| Intel Arc Pro B70 | llama.cpp | UD-Q8_K_XL | 115 tok/s @ 32,768 | candidate | localmaxxing |
| Radeon PRO W7800 | llama.cpp | UD-Q4_K_XL | 113 tok/s @ 8,192 | validated | 0xsero |
| GeForce RTX 5060 Ti | llama.cpp | Q3_K_M | 106 tok/s @ 2,048 | candidate | localmaxxing |
| Apple M5 Max 128GB | mlx | MLX 4bit | 102 tok/s @ 8,256 | candidate | exo-postgres |
| GeForce RTX 5060 Ti | llama.cpp | Q3_K_XL | 101 tok/s @ 2,048 | candidate | localmaxxing |
| Radeon AI PRO R9700 | llama.cpp | Q4_K_XL | 98.0 tok/s @ 2,048 | candidate | localmaxxing |
| Apple M4 Max 128GB | mlx | MLX 4bit | 92.3 tok/s @ 8,256 | candidate | exo-postgres |
From local-ai-registry (MIT), commit 124959d: 58 runs across 23 machines, 8 validated — model revision and runtime pinned, launch accepted — and 50 candidate, their label for useful evidence without a reproducible-launch promise. 31 more not listed.
| Hardware | Build measured | Decode @ 8K | Memory @ 8K | Decode mode |
|---|---|---|---|---|
| NVIDIA GeForce RTX 5090 32GB | NVFP4 | 147 tok/s | 29.4 GB | ordinary |
| NVIDIA RTX PRO 6000 Blackwell 96GB | NVFP4 | 139 tok/s | 89.8 GB | ordinary |
| NVIDIA RTX 6000 Ada 48GB | NVFP4 | 116 tok/s | 44.2 GB | ordinary |
| NVIDIA GeForce RTX 4090 24GB | NVFP4 | 114 tok/s | 21.6 GB | ordinary |
| NVIDIA DGX Spark 128GB | NVFP4 | 27.7 tok/s | 111.3 GB | ordinary |
| NVIDIA GeForce RTX 5090 32GB | NVFP4 | 247 tok/s | 29.3 GB | not stated |
| NVIDIA GeForce RTX 5090 32GB | NVFP4 | 247 tok/s | 29.3 GB | not stated |
| NVIDIA RTX PRO 6000 Blackwell 96GB | NVFP4 | 246 tok/s | 90.8 GB | not stated |
Measured by local.ai, not by us: 41 runs, 33 more not listed. Each names its engine pinned by image digest and the exact serve command, in the full record.
Registry record as of 2026-09-26. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.