architectures · deepseek-ai
Group Relative Policy Optimization (GRPO) on top of DeepSeek-V3 MLA base. No SFT needed — pure RL induces chain-of-thought thinking.
| Weight file SHA-256 | 4cd8fde22db0b147d8d4f542a00ce4bb5f251e7a36d0c2ae5cab4cef056e446a |
|---|
| Publisher | deepseek-ai |
|---|---|
| Origin | China |
| Licence | MIT |
| Modality | Text |
| Pulls recorded here | 88,001 |
Registry record as of 2026-06-06. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.