models · deepseek-ai
Llama-3-70B distilled from DeepSeek-R1 reasoning traces. Strong math and code at 70B scale.
| Weight file SHA-256 | e15541ba94dd226540a47bdec8ba43766e1144cc7007289b42712774363ff053 |
|---|---|
| Method | hf-published-sha256 |
| Publisher | deepseek-ai |
|---|---|
| Origin | China |
| Licence | MIT |
| Parameters | 70,553,706,496 (70.6B) |
| Context window | 8,192 tokens |
| Modality | text->text |
| Architecture | llama |
| Layers | 80 |
| Hidden size | 8192 |
| Attention | GQA (64 heads, 8 KV) |
| Native precision | BF16 |
| Pulls recorded here | 84,017 |
| Downloads reported upstream | 78,355 |
| MMLU-Pro | 79.5 |
|---|---|
| GPQA Diamond | 40.2 |
| AIME 2025 | 53.7 |
| MATH-500 | 93.5 |
| LiveCodeBench | 26.6 |
| IFBench | 27.6 |
| Humanity's Last Exam | 5.1 |
Reported by a third party and recorded here, not re-run by us. Source: artificialanalysis.ai/api/v2.
Registry record as of 2026-06-11. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.