models · google

Gemma 4 26B-A4B

Google Gemma 4 MoE variant. 26B total / 4B active parameters. 262K context. Efficient inference with 115 tok/s throughput.

Compare Verify

What this registry has checked

Quarantine. We hold the publisher’s own published hash for these weights and have recorded it against this entry. We have not downloaded and hashed the weights ourselves, so the hash is the publisher’s claim, not our measurement.
Weight file SHA-25625f5f6945c06b525d613f31a32c53d449de0e004f03d75ab5ddeb7a4d5823a48
Methodhf-published-sha256

What is recorded

Publishergoogle
OriginUnited States
LicenceApache-2.0
Parameters25,805,936,206 (26B-A4B)
Context window262K
Modalitytext+image+video->text
Native precisionBF16
Pulls recorded here180,015
Downloads reported upstream48,151

Reported scores

GPQA Diamond79.2
IFBench72.4
Humanity's Last Exam19.3

Reported by a third party and recorded here, not re-run by us. Source: artificialanalysis.ai/api/v2.

Measured on real hardware

None of these runs are ours. Each figure is the reporter’s, linked to their own evidence, under their own trust label. The number shown is single-stream decode — one user, concurrency 1 — because aggregate throughput across many users is a different quantity and reads as far faster than anyone will see.

Reported to local-ai-registry

HardwareEngineBuild measuredDecode, 1 streamLabelReported by
GeForce RTX 5090llama.cppQ3_K_M256 tok/s @ 2,048candidatelocalmaxxing
Intel Arc Pro B70llama.cppQ8_K_XL125 tok/s @ 32,768candidatelocalmaxxing
Radeon PRO W7800llama.cppQ4_K_S117 tok/s @ 8,192validated0xsero
Intel Arc Pro B70llama.cppUD-Q8_K_XL115 tok/s @ 32,768candidatelocalmaxxing
Radeon PRO W7800llama.cppUD-Q4_K_XL113 tok/s @ 8,192validated0xsero
GeForce RTX 5060 Tillama.cppQ3_K_M106 tok/s @ 2,048candidatelocalmaxxing
Apple M5 Max 128GBmlxMLX 4bit102 tok/s @ 8,256candidateexo-postgres
GeForce RTX 5060 Tillama.cppQ3_K_XL101 tok/s @ 2,048candidatelocalmaxxing
Radeon AI PRO R9700llama.cppQ4_K_XL98.0 tok/s @ 2,048candidatelocalmaxxing
Apple M4 Max 128GBmlxMLX 4bit92.3 tok/s @ 8,256candidateexo-postgres

From local-ai-registry (MIT), commit 124959d: 58 runs across 23 machines, 8 validated — model revision and runtime pinned, launch accepted — and 50 candidate, their label for useful evidence without a reproducible-launch promise. 31 more not listed.

Measured by local.ai

HardwareBuild measuredDecode @ 8KMemory @ 8KDecode mode
NVIDIA GeForce RTX 5090 32GBNVFP4147 tok/s29.4 GBordinary
NVIDIA RTX PRO 6000 Blackwell 96GBNVFP4139 tok/s89.8 GBordinary
NVIDIA RTX 6000 Ada 48GBNVFP4116 tok/s44.2 GBordinary
NVIDIA GeForce RTX 4090 24GBNVFP4114 tok/s21.6 GBordinary
NVIDIA DGX Spark 128GBNVFP427.7 tok/s111.3 GBordinary
NVIDIA GeForce RTX 5090 32GBNVFP4247 tok/s29.3 GBnot stated
NVIDIA GeForce RTX 5090 32GBNVFP4247 tok/s29.3 GBnot stated
NVIDIA RTX PRO 6000 Blackwell 96GBNVFP4246 tok/s90.8 GBnot stated

Measured by local.ai, not by us: 41 runs, 33 more not listed. Each names its engine pinned by image digest and the exact serve command, in the full record.

Elsewhere in the registry

Registry record as of 2026-09-26. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.