models · qwen

Qwen3.6-27B

Qwen3.6 dense 27.8B. 262K context, 150 tok/s. Arena 1138. GPQA Diamond 87.8%, LiveCodeBench 75.8%. Strong dense model at 28B scale.

Compare Verify

What this registry has checked

Quarantine. We hold the publisher’s own published hash for these weights and have recorded it against this entry. We have not downloaded and hashed the weights ourselves, so the hash is the publisher’s claim, not our measurement.
Weight file SHA-25620e9d0e734a8fb4e4f9a2ddd9aead33317e737c2b751bc601b6bec0173ab5237
Methodhf-published-sha256

What is recorded

Publisherqwen
OriginRecorded as a covered nation under 10 U.S.C. § 4872(f) — the PRC, Russia, Iran or North Korea. The registry did not record which.
LicenceApache-2.0
Parameters27,781,427,952 (27.8B)
Context window262K
Modalitytext+image+video->text
Native precisionBF16
Pulls recorded here420,014
Downloads reported upstream4,048,692

Reported scores

GPQA Diamond84.2
IFBench67.6
Humanity's Last Exam23.1

Reported by a third party and recorded here, not re-run by us. Source: artificialanalysis.ai/api/v2.

Measured on real hardware

None of these runs are ours. Each figure is the reporter’s, linked to their own evidence, under their own trust label. The number shown is single-stream decode — one user, concurrency 1 — because aggregate throughput across many users is a different quantity and reads as far faster than anyone will see.

Reported to local-ai-registry

HardwareEngineBuild measuredDecode, 1 streamLabelReported by
Radeon AI PRO R9700hipfireMQ4278 tok/s @ 2,048candidatelocalmaxxing
Radeon RX 7900 XTXhipfireMQ4252 tok/s @ 2,048candidatelocalmaxxing
GeForce RTX 3090 Tillama.cppUD-Q4_K_XL156 tok/s @ 4,096candidatelocalmaxxing
GeForce RTX 3090llama.cppUD-Q4_K_XL150 tok/s @ 2,048candidatelocalmaxxing
RTX PRO 6000 Blackwellllama.cppQ4_0144 tok/s @ 4,445candidatelocalmaxxing
RTX PRO 6000 BlackwellvllmNVFP4140 tok/s @ 2,048candidatelocalmaxxing
Ryzen AI Max+ 395hipfireMQ4112 tok/s @ 2,048candidatelocalmaxxing
GeForce RTX 5090 ×2 devicesvllmsafetensors NVFP485.3 tok/s @ 8,192candidatelocalmaxxing
GeForce RTX 5090llama.cppsafetensors IQ4_XS79.2 tok/s @ 8,192candidatelocalmaxxing
GeForce RTX 3090 ×2 devicesvllmsafetensors FP877.8 tok/s @ 8,192candidatelocalmaxxing

From local-ai-registry (MIT), commit 124959d: 306 runs across 32 machines, 10 validated — model revision and runtime pinned, launch accepted — and 296 candidate, their label for useful evidence without a reproducible-launch promise. 198 more not listed. 3 report figures that contradict each other — prefill below decode, which parallel prefill essentially never is — and are left to the source rather than ranked here.

Measured by local.ai

HardwareBuild measuredDecode @ 8KMemory @ 8KDecode mode
NVIDIA RTX PRO 6000 Blackwell 96GBQ4_K_M74.4 tok/s19.5 GBordinary
NVIDIA GeForce RTX 5090 32GBNVFP473.9 tok/s28.2 GBordinary
NVIDIA RTX PRO 6000 Blackwell 96GBNVFP473.0 tok/s90.2 GBordinary
NVIDIA GeForce RTX 5090 32GBNVFP465.8 tok/s28.3 GBordinary
NVIDIA RTX PRO 6000 Blackwell 96GBNVFP464.7 tok/s92.0 GBordinary
NVIDIA RTX 6000 Ada 48GBQ4_K_M44.5 tok/s22.7 GBordinary
NVIDIA RTX 6000 Ada 48GBNVFP438.0 tok/s43.4 GBordinary
Mac Studio M3 Ultra 96GB · 60-core GPU4bit28.9 tok/s23.2 GBordinary

Measured by local.ai, not by us: 112 runs, 104 more not listed. Each names its engine pinned by image digest and the exact serve command, in the full record.

Elsewhere in the registry

Registry record as of 2026-06-11. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.