models · qwen

Qwen3-14B

Qwen3 14B dense instruct model. Efficient mid-range option with strong reasoning and multilingual performance.

Compare Verify

What this registry has checked

Quarantine. We hold the publisher’s own published hash for these weights and have recorded it against this entry. We have not downloaded and hashed the weights ourselves, so the hash is the publisher’s claim, not our measurement.
Weight file SHA-256245877f20829ecc834373c6fc1e86dfb6bc6a699684dbb2967033d4f7a4fc786
Methodhf-published-sha256

What is recorded

Publisherqwen
OriginRecorded as a covered nation under 10 U.S.C. § 4872(f) — the PRC, Russia, Iran or North Korea. The registry did not record which.
LicenceApache-2.0
Parameters14,768,307,200 (14.8B)
Context window128K
Modalitytext->text
Architectureqwen3
Layers40
Hidden size5120
AttentionGQA (40 heads, 8 KV)
Native precisionBF16
Pulls recorded here750,019
Downloads reported upstream1,584,558

Reported scores

MMLU-Pro67.5
GPQA Diamond47
AIME 202558
MATH-50087.1
LiveCodeBench28
IFBench23.9
Humanity's Last Exam4.1

Reported by a third party and recorded here, not re-run by us. Source: artificialanalysis.ai/api/v2.

Measured on real hardware

None of these runs are ours. Each figure is the reporter’s, linked to their own evidence, under their own trust label. The number shown is single-stream decode — one user, concurrency 1 — because aggregate throughput across many users is a different quantity and reads as far faster than anyone will see.

Reported to local-ai-registry

HardwareEngineBuild measuredDecode, 1 streamLabelReported by
RTX PRO 6000 Blackwellvllmsafetensors FP890.8 tok/s @ 32,768candidatelocalmaxxing
GeForce RTX 3080llama.cppQ4_K_M75.0 tok/s @ 2,048candidatelocalmaxxing
GeForce RTX 5070llama.cppQ4_K_M66.6 tok/s @ 2,048candidatelocalmaxxing
GeForce RTX 3060llama.cppQ4_K_M37.0 tok/s @ 2,048candidatelocalmaxxing
GeForce RTX 3060ollamaQ4_K_M35.0 tok/s @ 2,048candidatelocalmaxxing
Apple M1 Max 64GBollamaQ4_K_M24.8 tok/s @ 2,048candidatelocalmaxxing
Measured outside 2K–32K context — decode speed depends heavily on context, so these are not comparable with the rows above
GeForce RTX 3090llama.cppQ4_K_M51.7 tok/s @ 1,024candidatelocalmaxxing
GeForce RTX 3090llama.cppQ4_K_M51.7 tok/s @ 512candidatelocalmaxxing
GeForce RTX 3060 ×2 devicesllama.cppQ4_K_M35.8 tok/s @ 512candidatelocalmaxxing
GeForce RTX 3060llama.cppQ4_K_M34.6 tok/s @ 512candidatelocalmaxxing

From local-ai-registry (MIT), commit 124959d: 10 runs across 6 machines, 0 validated — model revision and runtime pinned, launch accepted — and 10 candidate, their label for useful evidence without a reproducible-launch promise.

Elsewhere in the registry

Registry record as of 2026-06-11. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.