models · google

gemma-4-E2B-it

Catalogued from llm-stats Open LLM Leaderboard at Quarantine — unverified, pending attestation.

Compare Verify

What this registry has checked

Quarantine. We hold the publisher’s own published hash for these weights and have recorded it against this entry. We have not downloaded and hashed the weights ourselves, so the hash is the publisher’s claim, not our measurement.
Methodhf-published-sha256

What is recorded

Publishergoogle
OriginUnited States
LicenceApache-2.0
Parameters5,123,178,051 (5.12B)
Native precisionBF16
Pulls recorded here6
Downloads reported upstream2,952,135

Measured on real hardware

None of these runs are ours. Each figure is the reporter’s, linked to their own evidence, under their own trust label. The number shown is single-stream decode — one user, concurrency 1 — because aggregate throughput across many users is a different quantity and reads as far faster than anyone will see.

Measured by local.ai

HardwareBuild measuredDecode @ 8KMemory @ 8KDecode mode
NVIDIA GeForce RTX 5090 32GBUD-Q4_K_XL338 tok/s4.3 GBordinary
NVIDIA RTX PRO 6000 Blackwell 96GBUD-Q4_K_XL301 tok/s3.4 GBordinary
NVIDIA GeForce RTX 5090 32GBUD-Q8_K_XL282 tok/s5.5 GBordinary
NVIDIA GeForce RTX 5090 32GBFP8-Dynamic263 tok/s25.5 GBordinary
NVIDIA GeForce RTX 5090 32GBfp8263 tok/s25.5 GBordinary
NVIDIA RTX PRO 6000 Blackwell 96GBUD-Q8_K_XL253 tok/s5.5 GBordinary
NVIDIA RTX PRO 6000 Blackwell 96GBFP8-Dynamic249 tok/s85.7 GBordinary
NVIDIA RTX PRO 6000 Blackwell 96GBfp8249 tok/s85.7 GBordinary

Measured by local.ai, not by us: 84 runs, 76 more not listed. Each names its engine pinned by image digest and the exact serve command, in the full record.

Elsewhere in the registry

Registry record as of 2026-09-26. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.