architectures · deepseek-ai

DeepSeek-R1 Chain-of-Thought Architecture

Group Relative Policy Optimization (GRPO) on top of DeepSeek-V3 MLA base. No SFT needed — pure RL induces chain-of-thought thinking.

What this registry has checked

Quarantine. We hold the publisher’s own published hash for these weights and have recorded it against this entry. We have not downloaded and hashed the weights ourselves, so the hash is the publisher’s claim, not our measurement.
Weight file SHA-2564cd8fde22db0b147d8d4f542a00ce4bb5f251e7a36d0c2ae5cab4cef056e446a

What is recorded

Publisherdeepseek-ai
OriginChina
LicenceMIT
ModalityText
Pulls recorded here88,001

Elsewhere in the registry

Registry record as of 2026-06-06. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.