models · Qwen

Qwen3.5-9B

Catalogued from llm-stats Open LLM Leaderboard at Quarantine — unverified, pending attestation.

Compare Verify

What this registry has checked

Quarantine. We hold the publisher’s own published hash for these weights and have recorded it against this entry. We have not downloaded and hashed the weights ourselves, so the hash is the publisher’s claim, not our measurement.
Methodhf-published-sha256

What is recorded

PublisherQwen
OriginRecorded as a covered nation under 10 U.S.C. § 4872(f) — the PRC, Russia, Iran or North Korea. The registry did not record which.
LicenceApache-2.0
Parameters9,653,104,368 (9.65B)
Context window262,144 tokens
Modalitytext+image+video->text
Native precisionBF16
Pulls recorded here6
Downloads reported upstream9,468,124

Reported scores

GPQA Diamond80.6
IFBench66.7
Humanity's Last Exam14.9

Reported by a third party and recorded here, not re-run by us. Source: artificialanalysis.ai/api/v2.

Measured on real hardware

None of these runs are ours. Each figure is the reporter’s, linked to their own evidence, under their own trust label. The number shown is single-stream decode — one user, concurrency 1 — because aggregate throughput across many users is a different quantity and reads as far faster than anyone will see.

Reported to local-ai-registry

HardwareEngineBuild measuredDecode, 1 streamLabelReported by
GeForce RTX 5090llama.cppQ6_K170 tok/s @ 2,048candidatelocalmaxxing
GeForce RTX 5070 Tillama.cppQ4_K_S124 tok/s @ 8,192candidatelocalmaxxing
GeForce RTX 3090 Tillama.cppUD-Q4_K_XL119 tok/s @ 8,192candidatelocalmaxxing
GeForce RTX 3090 ×2 devicesllama.cppQ4_K_S110 tok/s @ 8,192candidatelocalmaxxing
H200 141GBollamaGGUF Q4_K_M103 tok/s @ 32,768candidate0xsero
Radeon RX 9070 XTllama.cppQ4_K_M87.5 tok/s @ 4,096candidatelocalmaxxing
Apple M3 Ultra 96GB 80-core GPUllama.cppGGUF Q4_K_M81.6 tok/s @ 8,256candidateexo-postgres
GeForce RTX 5080llama.cppQ8_079.6 tok/s @ 4,096candidatelocalmaxxing
Apple M5 Max 128GBllama.cppQ4_K_M79.3 tok/s @ 2,048candidatelocalmaxxing
Apple M5 Max 128GBllama.cppGGUF Q4_K_M78.6 tok/s @ 8,256candidateexo-postgres

From local-ai-registry (MIT), commit 124959d: 99 runs across 51 machines, 19 validated — model revision and runtime pinned, launch accepted — and 80 candidate, their label for useful evidence without a reproducible-launch promise. 48 more not listed. 21 report figures that contradict each other — prefill below decode, which parallel prefill essentially never is — and are left to the source rather than ranked here.

Measured by local.ai

HardwareBuild measuredDecode @ 8KMemory @ 8KDecode mode
NVIDIA GeForce RTX 5090 32GBQ4_K_M212 tok/s10.3 GBordinary
NVIDIA RTX PRO 6000 Blackwell 96GBQ4_K_M204 tok/s6.9 GBordinary
NVIDIA RTX PRO 6000 Blackwell 96GBFP8140 tok/s90.0 GBordinary
NVIDIA GeForce RTX 5090 32GBFP8139 tok/s28.4 GBordinary
NVIDIA GeForce RTX 4090 24GBQ4_K_M129 tok/s8.0 GBordinary
NVIDIA RTX 6000 Ada 48GBQ4_K_M127 tok/s14.6 GBordinary
Mac Studio M3 Ultra 96GB · 80-core GPU4bit94.0 tok/s7.5 GBordinary
Mac Studio M3 Ultra 96GB · 60-core GPU4bit90.9 tok/s7.5 GBordinary

Measured by local.ai, not by us: 133 runs, 125 more not listed. Each names its engine pinned by image digest and the exact serve command, in the full record.

Elsewhere in the registry

Registry record as of 2026-09-26. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.