models · Qwen

Qwen3-Coder-30B-A3B-Instruct

Catalogued from llm-stats Open LLM Leaderboard at Quarantine — unverified, pending attestation.

Compare Verify

What this registry has checked

Quarantine. We hold the publisher’s own published hash for these weights and have recorded it against this entry. We have not downloaded and hashed the weights ourselves, so the hash is the publisher’s claim, not our measurement.
Methodhf-published-sha256

What is recorded

PublisherQwen
OriginRecorded as a covered nation under 10 U.S.C. § 4872(f) — the PRC, Russia, Iran or North Korea. The registry did not record which.
LicenceApache-2.0
Parameters30,532,122,624 (30.5B)
Context window262,144 tokens
Modalitytext->text
Architectureqwen3_moe
Layers48
Hidden size2048
AttentionGQA (32 heads, 4 KV)
Native precisionBF16
Pulls recorded here6
Downloads reported upstream604,411

Reported scores

MMLU-Pro70.6
GPQA Diamond51.6
AIME 202529
MATH-50089.3
LiveCodeBench40.3
IFBench32.7
Humanity's Last Exam3.8

Reported by a third party and recorded here, not re-run by us. Source: artificialanalysis.ai/api/v2.

Measured on real hardware

None of these runs are ours. Each figure is the reporter’s, linked to their own evidence, under their own trust label. The number shown is single-stream decode — one user, concurrency 1 — because aggregate throughput across many users is a different quantity and reads as far faster than anyone will see.

Reported to local-ai-registry

HardwareEngineBuild measuredDecode, 1 streamLabelReported by
Intel Arc Pro B70llama.cppGGUF UD-Q4_K_XL108 tok/s @ 4,096candidatelocalmaxxing
Apple M3 Ultra 512GBllama.cppQ4_K_M100 tok/s @ 8,192candidatelocalmaxxing
Intel Arc Pro B70 ×2 devicesllama.cppUD-Q4_K_XL93.3 tok/s @ 2,048candidatelocalmaxxing
Ryzen AI Max+ 395llama.cppQ5_K_M80.5 tok/s @ 16,384candidatelocalmaxxing
Apple M1 Max 64GBollamaQ4_K_M72.3 tok/s @ 2,048candidatelocalmaxxing
Ryzen AI Max+ 395llama.cppQ4_K_M70.5 tok/s @ 4,096candidatelocalmaxxing
Radeon RX 7900 XTXllama.cppQ4_K_M43.8 tok/s @ 2,048candidatelocalmaxxing
Measured outside 2K–32K context — decode speed depends heavily on context, so these are not comparable with the rows above
Radeon PRO W7800llama.cppQ4_K_M109 tok/s @ 512candidate0xsero
Radeon RX 7900 XTXllama.cppQ4_K_M46.9 tok/s @ 1,024candidatelocalmaxxing
Radeon PRO W7800 ×2 devicesllama.cppQ4_K_M35.6 tok/s @ 512candidate0xsero

From local-ai-registry (MIT), commit 124959d: 14 runs across 9 machines, 0 validated — model revision and runtime pinned, launch accepted — and 14 candidate, their label for useful evidence without a reproducible-launch promise.

Elsewhere in the registry

Registry record as of 2026-09-26. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.