models · google
image-text-to-text model harvested from Hugging Face (trending). Catalogued at Quarantine — unverified, pending attestation.
| Method | hf-published-sha256 |
|---|
| Publisher | |
|---|---|
| Origin | United States |
| Licence | apache-2.0 |
| Parameters | 31,273,088,876 (31.3B) |
| Context window | 262,144 tokens |
| Modality | text+image+video->text |
| Native precision | BF16 |
| Pulls recorded here | 9,794,803 |
| Downloads reported upstream | 8,818,960 |
| Hardware | Engine | Build measured | Decode, 1 stream | Label | Reported by |
|---|---|---|---|---|---|
| GeForce RTX 5090 | llama.cpp | UD-Q4_K_XL-QAT | 68.5 tok/s @ 8,192 | candidate | localmaxxing |
| GeForce RTX 5090 | llama.cpp | Q6_K_XL | 51.7 tok/s @ 2,048 | candidate | localmaxxing |
| GeForce RTX 5090 | llama.cpp | UD-Q6_K_XL | 46.5 tok/s @ 8,192 | candidate | localmaxxing |
| GeForce RTX 3090 | llama.cpp | GGUF Q4_K_M | 38.9 tok/s @ 2,048 | candidate | localmaxxing |
| GeForce RTX 5090 | llama.cpp | Q8_0 | 38.0 tok/s @ 2,048 | candidate | localmaxxing |
| Radeon AI PRO R9700 | llama.cpp | Q4_K_M | 28.8 tok/s @ 2,048 | candidate | localmaxxing |
| Apple M3 Ultra 96GB 60-core GPU | mlx | MLX 4bit | 27.6 tok/s @ 8,256 | candidate | exo-postgres |
| Radeon AI PRO R9700 | llama.cpp | Q4_K_XL | 27.6 tok/s @ 2,048 | candidate | localmaxxing |
| Apple M3 Ultra 96GB 80-core GPU | mlx | MLX 4bit | 27.1 tok/s @ 8,256 | candidate | exo-postgres |
| Apple M5 Max 128GB | llama.cpp | UD-Q4_K_XL-QAT | 26.3 tok/s @ 2,048 | candidate | localmaxxing |
From local-ai-registry (MIT), commit 124959d: 57 runs across 16 machines, 0 validated — model revision and runtime pinned, launch accepted — and 57 candidate, their label for useful evidence without a reproducible-launch promise. 38 more not listed.
| Hardware | Build measured | Decode @ 8K | Memory @ 8K | Decode mode |
|---|---|---|---|---|
| NVIDIA RTX PRO 6000 Blackwell 96GB | Q4_K_M | 62.8 tok/s | 20.8 GB | ordinary |
| NVIDIA GeForce RTX 5090 32GB | AWQ-4bit | 61.9 tok/s | 28.2 GB | ordinary |
| NVIDIA RTX PRO 6000 Blackwell 96GB | AWQ-4bit | 59.6 tok/s | 89.6 GB | ordinary |
| NVIDIA RTX 6000 Ada 48GB | AWQ-4bit | 38.4 tok/s | 43.2 GB | ordinary |
| NVIDIA RTX 6000 Ada 48GB | Q4_K_M | 37.9 tok/s | 39.3 GB | ordinary |
| NVIDIA RTX PRO 6000 Blackwell 96GB | FP8 | 37.1 tok/s | 90.9 GB | ordinary |
| Mac Studio M3 Ultra 96GB · 60-core GPU | Q4_K_M | 26.7 tok/s | 32.6 GB | ordinary |
| Mac Studio M3 Ultra 96GB · 80-core GPU | Q4_K_M | 26.5 tok/s | 47.0 GB | ordinary |
Measured by local.ai, not by us: 67 runs, 59 more not listed. Each names its engine pinned by image digest and the exact serve command, in the full record.
Registry record as of 2026-09-01. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.