models · xiaomi
Xiaomi MiMo-V2-Flash. 309B. Arena 793. GPQA Diamond 83.7%, AIME 94.1%. Efficient flash-inference variant of the MiMo-V2 series.
| Weight file SHA-256 | 374e63be8f92083e357da700bf29ee0c1ad708347595713c4ffe04bb3c1697b6 |
|---|---|
| Method | hf-published-sha256 |
| Publisher | xiaomi |
|---|---|
| Origin | Recorded as a covered nation under 10 U.S.C. § 4872(f) — the PRC, Russia, Iran or North Korea. The registry did not record which. |
| Licence | Apache-2.0 |
| Publisher’s own licence statement | mit |
| Parameters | 8,306,217,216 (8.31B) |
| Context window | 128K |
| Architecture | qwen2_5_vl |
| Layers | 36 |
| Hidden size | 4096 |
| Attention | GQA (32 heads, 8 KV) |
| Native precision | BF16 |
| Pulls recorded here | 160,012 |
| Downloads reported upstream | 1,128 |
| GPQA Diamond | 83.5 |
|---|---|
| IFBench | 71.8 |
| Humanity's Last Exam | 22.1 |
Reported by a third party and recorded here, not re-run by us. Source: artificialanalysis.ai/api/v2.
| Hardware | Build measured | Decode @ 8K | Memory @ 8K | Decode mode |
|---|---|---|---|---|
| MacBook Pro M5 Max 128GB · 40-core GPU | UD-IQ2_M | 57.2 tok/s | 111.7 GB | ordinary |
| MacBook Pro M5 Max 128GB · 40-core GPU | UD-IQ2_XXS | 56.1 tok/s | 106.6 GB | ordinary |
| Mac Studio M3 Ultra 96GB · 80-core GPU | UD-IQ2_XXS | 50.7 tok/s | 51.2 GB | ordinary |
| Mac Studio M3 Ultra 96GB · 80-core GPU | UD-IQ2_M | 50.4 tok/s | 51.2 GB | ordinary |
| MacBook Pro M4 Max 128GB · 40-core GPU | UD-IQ2_M | 39.0 tok/s | 111.1 GB | ordinary |
| NVIDIA DGX Spark 128GB | UD-IQ2_XXS | 24.8 tok/s | 108.0 GB | ordinary |
| NVIDIA DGX Spark 128GB | UD-IQ2_M | 24.3 tok/s | 113.1 GB | ordinary |
Measured by local.ai, not by us: 7 runs. Each names its engine pinned by image digest and the exact serve command, in the full record.
Registry record as of 2026-06-11. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.