models · google
Catalogued from local.ai local-inference benchmark index at Quarantine — unverified, pending attestation.
| Method | hf-published-sha256 |
|---|
| Publisher | |
|---|---|
| Origin | United States |
| Licence | Apache-2.0 |
| Parameters | 25,805,936,206 (25.8B) |
| Context window | 262,144 tokens |
| Modality | text+image+video->text |
| Native precision | BF16 |
| Downloads reported upstream | 8,818,645 |
| Hardware | Engine | Build measured | Decode, 1 stream | Label | Reported by |
|---|---|---|---|---|---|
| GeForce RTX 5090 | llama.cpp | Q3_K_M | 256 tok/s @ 2,048 | candidate | localmaxxing |
| Intel Arc Pro B70 | llama.cpp | Q8_K_XL | 125 tok/s @ 32,768 | candidate | localmaxxing |
| Radeon PRO W7800 | llama.cpp | Q4_K_S | 117 tok/s @ 8,192 | validated | 0xsero |
| Intel Arc Pro B70 | llama.cpp | UD-Q8_K_XL | 115 tok/s @ 32,768 | candidate | localmaxxing |
| Radeon PRO W7800 | llama.cpp | UD-Q4_K_XL | 113 tok/s @ 8,192 | validated | 0xsero |
| GeForce RTX 5060 Ti | llama.cpp | Q3_K_M | 106 tok/s @ 2,048 | candidate | localmaxxing |
| Apple M5 Max 128GB | mlx | MLX 4bit | 102 tok/s @ 8,256 | candidate | exo-postgres |
| GeForce RTX 5060 Ti | llama.cpp | Q3_K_XL | 101 tok/s @ 2,048 | candidate | localmaxxing |
| Radeon AI PRO R9700 | llama.cpp | Q4_K_XL | 98.0 tok/s @ 2,048 | candidate | localmaxxing |
| Apple M4 Max 128GB | mlx | MLX 4bit | 92.3 tok/s @ 8,256 | candidate | exo-postgres |
From local-ai-registry (MIT), commit 124959d: 58 runs across 23 machines, 8 validated — model revision and runtime pinned, launch accepted — and 50 candidate, their label for useful evidence without a reproducible-launch promise. 31 more not listed.
| Hardware | Build measured | Decode @ 8K | Memory @ 8K | Decode mode |
|---|---|---|---|---|
| NVIDIA RTX PRO 6000 Blackwell 96GB | AWQ-4bit | 202 tok/s | 90.4 GB | ordinary |
| NVIDIA GeForce RTX 5090 32GB | AWQ-4bit | 192 tok/s | 29.3 GB | ordinary |
| NVIDIA RTX PRO 6000 Blackwell 96GB | fp8 | 165 tok/s | 90.7 GB | ordinary |
| NVIDIA RTX 6000 Ada 48GB | AWQ-4bit | 146 tok/s | 44.4 GB | ordinary |
| NVIDIA GeForce RTX 4090 24GB | AWQ-4bit | 132 tok/s | 21.8 GB | ordinary |
| NVIDIA RTX 6000 Ada 48GB | fp8 | 119 tok/s | 44.3 GB | ordinary |
| MacBook Pro M5 Max 128GB · 40-core GPU | 4bit | 103 tok/s | 16.6 GB | ordinary |
| MacBook Pro M4 Max 128GB · 40-core GPU | 4bit | 92.5 tok/s | 28.5 GB | ordinary |
Measured by local.ai, not by us: 67 runs, 59 more not listed. Each names its engine pinned by image digest and the exact serve command, in the full record.
Registry record as of 2026-09-26. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.