models · google

gemma-4-31B-it

image-text-to-text model harvested from Hugging Face (trending). Catalogued at Quarantine — unverified, pending attestation.

Compare Verify

What this registry has checked

Quarantine. We hold the publisher’s own published hash for these weights and have recorded it against this entry. We have not downloaded and hashed the weights ourselves, so the hash is the publisher’s claim, not our measurement.
Methodhf-published-sha256

What is recorded

Publishergoogle
OriginUnited States
Licenceapache-2.0
Parameters31,273,088,876 (31.3B)
Context window262,144 tokens
Modalitytext+image+video->text
Native precisionBF16
Pulls recorded here9,794,803
Downloads reported upstream8,818,960

Measured on real hardware

None of these runs are ours. Each figure is the reporter’s, linked to their own evidence, under their own trust label. The number shown is single-stream decode — one user, concurrency 1 — because aggregate throughput across many users is a different quantity and reads as far faster than anyone will see.

Reported to local-ai-registry

HardwareEngineBuild measuredDecode, 1 streamLabelReported by
GeForce RTX 5090llama.cppUD-Q4_K_XL-QAT68.5 tok/s @ 8,192candidatelocalmaxxing
GeForce RTX 5090llama.cppQ6_K_XL51.7 tok/s @ 2,048candidatelocalmaxxing
GeForce RTX 5090llama.cppUD-Q6_K_XL46.5 tok/s @ 8,192candidatelocalmaxxing
GeForce RTX 3090llama.cppGGUF Q4_K_M38.9 tok/s @ 2,048candidatelocalmaxxing
GeForce RTX 5090llama.cppQ8_038.0 tok/s @ 2,048candidatelocalmaxxing
Radeon AI PRO R9700llama.cppQ4_K_M28.8 tok/s @ 2,048candidatelocalmaxxing
Apple M3 Ultra 96GB 60-core GPUmlxMLX 4bit27.6 tok/s @ 8,256candidateexo-postgres
Radeon AI PRO R9700llama.cppQ4_K_XL27.6 tok/s @ 2,048candidatelocalmaxxing
Apple M3 Ultra 96GB 80-core GPUmlxMLX 4bit27.1 tok/s @ 8,256candidateexo-postgres
Apple M5 Max 128GBllama.cppUD-Q4_K_XL-QAT26.3 tok/s @ 2,048candidatelocalmaxxing

From local-ai-registry (MIT), commit 124959d: 57 runs across 16 machines, 0 validated — model revision and runtime pinned, launch accepted — and 57 candidate, their label for useful evidence without a reproducible-launch promise. 38 more not listed.

Measured by local.ai

HardwareBuild measuredDecode @ 8KMemory @ 8KDecode mode
NVIDIA RTX PRO 6000 Blackwell 96GBQ4_K_M62.8 tok/s20.8 GBordinary
NVIDIA GeForce RTX 5090 32GBAWQ-4bit61.9 tok/s28.2 GBordinary
NVIDIA RTX PRO 6000 Blackwell 96GBAWQ-4bit59.6 tok/s89.6 GBordinary
NVIDIA RTX 6000 Ada 48GBAWQ-4bit38.4 tok/s43.2 GBordinary
NVIDIA RTX 6000 Ada 48GBQ4_K_M37.9 tok/s39.3 GBordinary
NVIDIA RTX PRO 6000 Blackwell 96GBFP837.1 tok/s90.9 GBordinary
Mac Studio M3 Ultra 96GB · 60-core GPUQ4_K_M26.7 tok/s32.6 GBordinary
Mac Studio M3 Ultra 96GB · 80-core GPUQ4_K_M26.5 tok/s47.0 GBordinary

Measured by local.ai, not by us: 67 runs, 59 more not listed. Each names its engine pinned by image digest and the exact serve command, in the full record.

Elsewhere in the registry

Registry record as of 2026-09-01. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.