models · openai
gpt-oss-20b — catalogued from third-party evaluation data; no weights held by Sovereign Frontier.
| Method | hf-published-sha256 |
|---|
| Publisher | openai |
|---|---|
| Origin | United States |
| Licence | Apache-2.0 |
| Parameters | 20,914,757,184 (20.9B) |
| Context window | 131,072 tokens |
| Modality | text->text |
| Architecture | gpt_oss |
| Layers | 24 |
| Hidden size | 2880 |
| Attention | GQA (64 heads, 8 KV) |
| Native precision | U8 |
| Pulls recorded here | 10 |
| Downloads reported upstream | 6,627,327 |
| MMLU-Pro | 74.8 |
|---|---|
| GPQA Diamond | 68.8 |
| AIME 2025 | 89.3 |
| LiveCodeBench | 77.7 |
| IFBench | 65.1 |
| Humanity's Last Exam | 11 |
Reported by a third party and recorded here, not re-run by us. Source: artificialanalysis.ai/api/v2.
| Hardware | Engine | Build measured | Decode, 1 stream | Label | Reported by |
|---|---|---|---|---|---|
| GeForce RTX 5080 | llama.cpp | Q8_0 | 222 tok/s @ 4,096 | candidate | localmaxxing |
| GeForce RTX 3060 | ollama | Q4_K_M | 74.0 tok/s @ 2,048 | candidate | localmaxxing |
| Apple M1 Max 64GB | ollama | MXFP4 | 62.4 tok/s @ 2,048 | candidate | localmaxxing |
| Intel Arc Pro B60 | llama.cpp | Q8_0 | 48.3 tok/s @ 4,096 | candidate | localmaxxing |
| GeForce RTX 3080 | ollama | MXFP4 | 21.0 tok/s @ 2,048 | candidate | localmaxxing |
| Measured outside 2K–32K context — decode speed depends heavily on context, so these are not comparable with the rows above | |||||
| GeForce RTX 3090 | llama.cpp | Q4_K_M | 180 tok/s @ 512 | candidate | localmaxxing |
| Apple M1 Max 64GB | llama.cpp | Q8_0 | 83.6 tok/s @ 640 | candidate | localmaxxing |
From local-ai-registry (MIT), commit 124959d: 11 runs across 10 machines, 0 validated — model revision and runtime pinned, launch accepted — and 11 candidate, their label for useful evidence without a reproducible-launch promise. 1 report figures that contradict each other — prefill below decode, which parallel prefill essentially never is — and are left to the source rather than ranked here.
| Hardware | Build measured | Decode @ 8K | Memory @ 8K | Decode mode |
|---|---|---|---|---|
| NVIDIA GeForce RTX 5090 32GB | — | 286 tok/s | 30.5 GB | ordinary |
| NVIDIA GeForce RTX 5090 32GB | mxfp4 | 286 tok/s | 30.5 GB | ordinary |
| NVIDIA RTX PRO 6000 Blackwell 96GB | — | 272 tok/s | 90.4 GB | ordinary |
| NVIDIA RTX PRO 6000 Blackwell 96GB | mxfp4 | 272 tok/s | 90.4 GB | ordinary |
| NVIDIA GeForce RTX 4090 24GB | — | 193 tok/s | 23.2 GB | ordinary |
| NVIDIA GeForce RTX 4090 24GB | mxfp4 | 193 tok/s | 23.2 GB | ordinary |
| NVIDIA RTX 6000 Ada 48GB | — | 181 tok/s | 45.8 GB | ordinary |
| NVIDIA RTX 6000 Ada 48GB | mxfp4 | 181 tok/s | 45.8 GB | ordinary |
Measured by local.ai, not by us: 28 runs, 20 more not listed. Each names its engine pinned by image digest and the exact serve command, in the full record.
Registry record as of 2026-09-26. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.