models · qwen
image-text-to-text model harvested from Hugging Face (trending). Catalogued at Quarantine — unverified, pending attestation.
| Method | hf-published-sha256 |
|---|
| Publisher | qwen |
|---|---|
| Origin | Recorded as a covered nation under 10 U.S.C. § 4872(f) — the PRC, Russia, Iran or North Korea. The registry did not record which. |
| Licence | apache-2.0 |
| Parameters | 35,951,822,704 (36B) |
| Context window | 262,144 tokens |
| Modality | text+image+video->text |
| Native precision | BF16 |
| Pulls recorded here | 5,129,809 |
| Downloads reported upstream | 3,654,078 |
| GPQA Diamond | 84.1 |
|---|---|
| IFBench | 64.4 |
| Humanity's Last Exam | 22.2 |
Reported by a third party and recorded here, not re-run by us. Source: artificialanalysis.ai/api/v2.
| Hardware | Engine | Build measured | Decode, 1 stream | Label | Reported by |
|---|---|---|---|---|---|
| Intel Arc Pro B70 | vllm | GPTQ-Int4 | 1140 tok/s @ 16,384 | candidate | localmaxxing |
| Radeon RX 7900 XTX | hipfire | MQ4R | 495 tok/s @ 4,096 | candidate | localmaxxing |
| GeForce RTX 5090 ×2 devices | vllm | safetensors NVFP4 | 341 tok/s @ 32,768 | candidate | localmaxxing |
| GeForce RTX 5090 | llama.cpp | Q5_K_M | 265 tok/s @ 2,048 | candidate | localmaxxing |
| GeForce RTX 5090 | llama.cpp | Unsloth-Dynamic-Q4_K_M | 262 tok/s @ 32,768 | candidate | localmaxxing |
| GeForce RTX 5090 | llama.cpp | safetensors Q4_K_M | 231 tok/s @ 4,096 | candidate | localmaxxing |
| GeForce RTX 5090 | llama.cpp | UD-Q6_K | 220 tok/s @ 4,096 | candidate | localmaxxing |
| GeForce RTX 4090 | llama.cpp | Q4_K_M | 214 tok/s @ 32,768 | candidate | localmaxxing |
| GeForce RTX 5080 | llama.cpp | APEX-MTP-I-Mini | 188 tok/s @ 4,096 | candidate | localmaxxing |
| GeForce RTX 5090 | ollama | safetensors Q4_K_M | 176 tok/s @ 32,768 | candidate | localmaxxing |
From local-ai-registry (MIT), commit 124959d: 384 runs across 48 machines, 12 validated — model revision and runtime pinned, launch accepted — and 372 candidate, their label for useful evidence without a reproducible-launch promise. 288 more not listed. 26 report figures that contradict each other — prefill below decode, which parallel prefill essentially never is — and are left to the source rather than ranked here.
| Hardware | Build measured | Decode @ 8K | Memory @ 8K | Decode mode |
|---|---|---|---|---|
| NVIDIA GeForce RTX 5090 32GB | iq1m | 268 tok/s | 13.9 GB | ordinary |
| NVIDIA GeForce RTX 5090 32GB | iq2xxs | 268 tok/s | 14.5 GB | ordinary |
| NVIDIA GeForce RTX 5090 32GB | UD-IQ2_M | 259 tok/s | 15.0 GB | ordinary |
| NVIDIA GeForce RTX 5090 32GB | UD-Q2_K_XL | 257 tok/s | 15.6 GB | ordinary |
| NVIDIA RTX PRO 6000 Blackwell 96GB | UD-Q2_K_XL | 255 tok/s | 18.5 GB | ordinary |
| NVIDIA RTX PRO 6000 Blackwell 96GB | UD-Q4_K_S | 251 tok/s | 26.9 GB | ordinary |
| NVIDIA RTX PRO 6000 Blackwell 96GB | q3km | 250 tok/s | 22.7 GB | ordinary |
| NVIDIA RTX PRO 6000 Blackwell 96GB | UD-Q3_K_XL | 250 tok/s | 22.8 GB | ordinary |
Measured by local.ai, not by us: 382 runs, 374 more not listed. Each names its engine pinned by image digest and the exact serve command, in the full record.
Registry record as of 2026-09-01. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.