models · qwen

Qwen3.6-35B-A3B

image-text-to-text model harvested from Hugging Face (trending). Catalogued at Quarantine — unverified, pending attestation.

Compare Verify

What this registry has checked

Quarantine. We hold the publisher’s own published hash for these weights and have recorded it against this entry. We have not downloaded and hashed the weights ourselves, so the hash is the publisher’s claim, not our measurement.
Methodhf-published-sha256

What is recorded

Publisherqwen
OriginRecorded as a covered nation under 10 U.S.C. § 4872(f) — the PRC, Russia, Iran or North Korea. The registry did not record which.
Licenceapache-2.0
Parameters35,951,822,704 (36B)
Context window262,144 tokens
Modalitytext+image+video->text
Native precisionBF16
Pulls recorded here5,129,809
Downloads reported upstream3,654,078

Reported scores

GPQA Diamond84.1
IFBench64.4
Humanity's Last Exam22.2

Reported by a third party and recorded here, not re-run by us. Source: artificialanalysis.ai/api/v2.

Measured on real hardware

None of these runs are ours. Each figure is the reporter’s, linked to their own evidence, under their own trust label. The number shown is single-stream decode — one user, concurrency 1 — because aggregate throughput across many users is a different quantity and reads as far faster than anyone will see.

Reported to local-ai-registry

HardwareEngineBuild measuredDecode, 1 streamLabelReported by
Intel Arc Pro B70vllmGPTQ-Int41140 tok/s @ 16,384candidatelocalmaxxing
Radeon RX 7900 XTXhipfireMQ4R495 tok/s @ 4,096candidatelocalmaxxing
GeForce RTX 5090 ×2 devicesvllmsafetensors NVFP4341 tok/s @ 32,768candidatelocalmaxxing
GeForce RTX 5090llama.cppQ5_K_M265 tok/s @ 2,048candidatelocalmaxxing
GeForce RTX 5090llama.cppUnsloth-Dynamic-Q4_K_M262 tok/s @ 32,768candidatelocalmaxxing
GeForce RTX 5090llama.cppsafetensors Q4_K_M231 tok/s @ 4,096candidatelocalmaxxing
GeForce RTX 5090llama.cppUD-Q6_K220 tok/s @ 4,096candidatelocalmaxxing
GeForce RTX 4090llama.cppQ4_K_M214 tok/s @ 32,768candidatelocalmaxxing
GeForce RTX 5080llama.cppAPEX-MTP-I-Mini188 tok/s @ 4,096candidatelocalmaxxing
GeForce RTX 5090ollamasafetensors Q4_K_M176 tok/s @ 32,768candidatelocalmaxxing

From local-ai-registry (MIT), commit 124959d: 384 runs across 48 machines, 12 validated — model revision and runtime pinned, launch accepted — and 372 candidate, their label for useful evidence without a reproducible-launch promise. 288 more not listed. 26 report figures that contradict each other — prefill below decode, which parallel prefill essentially never is — and are left to the source rather than ranked here.

Measured by local.ai

HardwareBuild measuredDecode @ 8KMemory @ 8KDecode mode
NVIDIA GeForce RTX 5090 32GBiq1m268 tok/s13.9 GBordinary
NVIDIA GeForce RTX 5090 32GBiq2xxs268 tok/s14.5 GBordinary
NVIDIA GeForce RTX 5090 32GBUD-IQ2_M259 tok/s15.0 GBordinary
NVIDIA GeForce RTX 5090 32GBUD-Q2_K_XL257 tok/s15.6 GBordinary
NVIDIA RTX PRO 6000 Blackwell 96GBUD-Q2_K_XL255 tok/s18.5 GBordinary
NVIDIA RTX PRO 6000 Blackwell 96GBUD-Q4_K_S251 tok/s26.9 GBordinary
NVIDIA RTX PRO 6000 Blackwell 96GBq3km250 tok/s22.7 GBordinary
NVIDIA RTX PRO 6000 Blackwell 96GBUD-Q3_K_XL250 tok/s22.8 GBordinary

Measured by local.ai, not by us: 382 runs, 374 more not listed. Each names its engine pinned by image digest and the exact serve command, in the full record.

Elsewhere in the registry

Registry record as of 2026-09-01. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.