models · qwen

Qwen3.5-35B-A3B

Qwen3.5 35B MoE with 3B active params. Ultra-fast inference at 233 tok/s.

Compare Verify

What this registry has checked

Quarantine. We hold the publisher’s own published hash for these weights and have recorded it against this entry. We have not downloaded and hashed the weights ourselves, so the hash is the publisher’s claim, not our measurement.
Weight file SHA-256e1dfb292d55fb73a30e536d1f4a3b73a2d56c151369b807ad1e0a11a321c880e
Methodhf-published-sha256

What is recorded

Publisherqwen
OriginChina
LicenceApache-2.0
Parameters35,951,822,704 (36B)
Context window262,144 tokens
Modalitytext+image+video->text
Native precisionBF16
Pulls recorded here130,757
Downloads reported upstream2,048,891

Reported scores

GPQA Diamond84.5
IFBench72.5
Humanity's Last Exam21

Reported by a third party and recorded here, not re-run by us. Source: artificialanalysis.ai/api/v2.

Measured on real hardware

None of these runs are ours. Each figure is the reporter’s, linked to their own evidence, under their own trust label. The number shown is single-stream decode — one user, concurrency 1 — because aggregate throughput across many users is a different quantity and reads as far faster than anyone will see.

Reported to local-ai-registry

HardwareEngineBuild measuredDecode, 1 streamLabelReported by
GeForce RTX 5070llama.cppIQ1_M182 tok/s @ 2,048candidatelocalmaxxing
GeForce RTX 4090llama.cppUD-Q4_K_XL157 tok/s @ 2,048candidatelocalmaxxing
GeForce RTX 3090llama.cppUD-Q4_K_XL124 tok/s @ 2,048candidatelocalmaxxing
Ryzen AI Max+ 395llama.cppQ4_K_M73.8 tok/s @ 32,768candidatelocalmaxxing
Ryzen AI Max+ 395llama.cppQ4_K_M58.3 tok/s @ 8,192candidatelocalmaxxing
Measured outside 2K–32K context — decode speed depends heavily on context, so these are not comparable with the rows above
Apple M5 Max 128GBollamaNVFP4122 tok/s @ 262,144candidatelocalmaxxing
Ryzen AI Max+ 395llama.cppQ4_K_M72.7 tok/s @ 80,000candidatelocalmaxxing
Intel Arc Pro B70llama.cppQ4_K_XL62.8 tok/s @ 262,144candidatelocalmaxxing
Intel Arc Pro B70llama.cppQ5_K_M61.5 tok/s @ 262,144candidatelocalmaxxing
NVIDIA DGX Spark GB10llama.cppUD-Q4_K_XL60.3 tok/s @ 262,144candidatelocalmaxxing

From local-ai-registry (MIT), commit 124959d: 21 runs across 12 machines, 0 validated — model revision and runtime pinned, launch accepted — and 21 candidate, their label for useful evidence without a reproducible-launch promise. 2 more not listed. 1 report figures that contradict each other — prefill below decode, which parallel prefill essentially never is — and are left to the source rather than ranked here.

Elsewhere in the registry

Registry record as of 2026-06-11. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.