models · deepseek-ai

DeepSeek-R1-Zero

DeepSeek-R1-Zero — RL-trained reasoning model without SFT warmup. Demonstrates emergent chain-of-thought.

Compare Verify

What this registry has checked

Quarantine. We hold the publisher’s own published hash for these weights and have recorded it against this entry. We have not downloaded and hashed the weights ourselves, so the hash is the publisher’s claim, not our measurement.
Weight file SHA-256bf4e7665d5edd7a80615864e108491b7a44cd9cdcb4134b5dac1d95f1ef8e63d
Methodhf-published-sha256

What is recorded

Publisherdeepseek-ai
OriginChina
LicenceMIT
Parameters684,531,386,000 (684.5B)
Architecturedeepseek_v3
Layers61
Hidden size7168
AttentionMLA (128 heads, 128 KV)
Native precisionF8_E4M3
Pulls recorded here76,019
Downloads reported upstream10,636

Reported scores

GPQA Diamond73.3
AIME 202571

Reported by a third party and recorded here, not re-run by us.

Elsewhere in the registry

Registry record as of 2026-06-11. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.