models · openai
OpenAI's 120B open-source model released under Apache-2.0. Strong general reasoning and coding.
| Weight file SHA-256 | fee26c4f4e06ca8f9dcb5b135eaeacb732eb10eaae91b5076e2b7fabdf4336e6 |
|---|---|
| Method | hf-published-sha256 |
| Publisher | openai |
|---|---|
| Origin | United States |
| Licence | Apache-2.0 |
| Parameters | 116,829,156,672 (116.8B) |
| Context window | 131,072 tokens |
| Modality | text->text |
| Architecture | gpt_oss |
| Layers | 36 |
| Hidden size | 2880 |
| Attention | GQA (64 heads, 8 KV) |
| Native precision | U8 |
| Pulls recorded here | 156,019 |
| Downloads reported upstream | 5,378,443 |
| MMLU-Pro | 80.8 |
|---|---|
| GPQA Diamond | 78.2 |
| AIME 2025 | 93.4 |
| LiveCodeBench | 87.8 |
| IFBench | 69 |
| Humanity's Last Exam | 19.6 |
Reported by a third party and recorded here, not re-run by us. Source: artificialanalysis.ai/api/v2.
| Hardware | Engine | Build measured | Decode, 1 stream | Label | Reported by |
|---|---|---|---|---|---|
| Intel Arc Pro B70 ×2 devices | llama.cpp | MXFP4 | 59.4 tok/s @ 2,048 | candidate | localmaxxing |
| NVIDIA DGX Spark GB10 | vllm | MXFP4 | 34.4 tok/s @ 32,768 | candidate | localmaxxing |
| Measured outside 2K–32K context — decode speed depends heavily on context, so these are not comparable with the rows above | |||||
| RTX PRO 6000 Blackwell | llama.cpp | F16 | 203 tok/s @ 845 | candidate | localmaxxing |
| Radeon AI PRO R9700 ×3 devices | llama.cpp | F16 | 72.8 tok/s @ 845 | candidate | localmaxxing |
| Ryzen AI Max+ 395 | llama.cpp | F16 | 38.8 tok/s @ 845 | candidate | localmaxxing |
From local-ai-registry (MIT), commit 124959d: 6 runs across 5 machines, 0 validated — model revision and runtime pinned, launch accepted — and 6 candidate, their label for useful evidence without a reproducible-launch promise.
| Hardware | Build measured | Decode @ 8K | Memory @ 8K | Decode mode |
|---|---|---|---|---|
| NVIDIA RTX PRO 6000 Blackwell 96GB | — | 137 tok/s | 90.9 GB | not stated |
| NVIDIA RTX PRO 6000 Blackwell 96GB | — | 128 tok/s | 94.1 GB | not stated |
| NVIDIA RTX PRO 6000 Blackwell 96GB | — | 124 tok/s | 94.1 GB | not stated |
| MacBook Pro M5 Max 128GB · 40-core GPU | MXFP4 | 71.3 tok/s | 64.8 GB | not stated |
| Mac Studio M3 Ultra 96GB · 80-core GPU | MXFP4 | 67.9 tok/s | 64.8 GB | not stated |
| MacBook Pro M4 Max 128GB · 40-core GPU | MXFP4 | 62.1 tok/s | 64.8 GB | not stated |
| NVIDIA DGX Spark 128GB | — | 33.7 tok/s | 102.8 GB | not stated |
| NVIDIA DGX Spark 128GB | — | 33.0 tok/s | 88.7 GB | not stated |
Measured by local.ai, not by us: 16 runs, 8 more not listed. Each names its engine pinned by image digest and the exact serve command, in the full record.
Registry record as of 2026-08-13. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.