models · openai

gpt-oss-20b

gpt-oss-20b — catalogued from third-party evaluation data; no weights held by Sovereign Frontier.

Compare Verify

What this registry has checked

Quarantine. We hold the publisher’s own published hash for these weights and have recorded it against this entry. We have not downloaded and hashed the weights ourselves, so the hash is the publisher’s claim, not our measurement.
Methodhf-published-sha256

What is recorded

Publisheropenai
OriginUnited States
LicenceApache-2.0
Parameters20,914,757,184 (20.9B)
Context window131,072 tokens
Modalitytext->text
Architecturegpt_oss
Layers24
Hidden size2880
AttentionGQA (64 heads, 8 KV)
Native precisionU8
Pulls recorded here10
Downloads reported upstream6,627,327

Reported scores

MMLU-Pro74.8
GPQA Diamond68.8
AIME 202589.3
LiveCodeBench77.7
IFBench65.1
Humanity's Last Exam11

Reported by a third party and recorded here, not re-run by us. Source: artificialanalysis.ai/api/v2.

Measured on real hardware

None of these runs are ours. Each figure is the reporter’s, linked to their own evidence, under their own trust label. The number shown is single-stream decode — one user, concurrency 1 — because aggregate throughput across many users is a different quantity and reads as far faster than anyone will see.

Reported to local-ai-registry

HardwareEngineBuild measuredDecode, 1 streamLabelReported by
GeForce RTX 5080llama.cppQ8_0222 tok/s @ 4,096candidatelocalmaxxing
GeForce RTX 3060ollamaQ4_K_M74.0 tok/s @ 2,048candidatelocalmaxxing
Apple M1 Max 64GBollamaMXFP462.4 tok/s @ 2,048candidatelocalmaxxing
Intel Arc Pro B60llama.cppQ8_048.3 tok/s @ 4,096candidatelocalmaxxing
GeForce RTX 3080ollamaMXFP421.0 tok/s @ 2,048candidatelocalmaxxing
Measured outside 2K–32K context — decode speed depends heavily on context, so these are not comparable with the rows above
GeForce RTX 3090llama.cppQ4_K_M180 tok/s @ 512candidatelocalmaxxing
Apple M1 Max 64GBllama.cppQ8_083.6 tok/s @ 640candidatelocalmaxxing

From local-ai-registry (MIT), commit 124959d: 11 runs across 10 machines, 0 validated — model revision and runtime pinned, launch accepted — and 11 candidate, their label for useful evidence without a reproducible-launch promise. 1 report figures that contradict each other — prefill below decode, which parallel prefill essentially never is — and are left to the source rather than ranked here.

Measured by local.ai

HardwareBuild measuredDecode @ 8KMemory @ 8KDecode mode
NVIDIA GeForce RTX 5090 32GB—286 tok/s30.5 GBordinary
NVIDIA GeForce RTX 5090 32GBmxfp4286 tok/s30.5 GBordinary
NVIDIA RTX PRO 6000 Blackwell 96GB—272 tok/s90.4 GBordinary
NVIDIA RTX PRO 6000 Blackwell 96GBmxfp4272 tok/s90.4 GBordinary
NVIDIA GeForce RTX 4090 24GB—193 tok/s23.2 GBordinary
NVIDIA GeForce RTX 4090 24GBmxfp4193 tok/s23.2 GBordinary
NVIDIA RTX 6000 Ada 48GB—181 tok/s45.8 GBordinary
NVIDIA RTX 6000 Ada 48GBmxfp4181 tok/s45.8 GBordinary

Measured by local.ai, not by us: 28 runs, 20 more not listed. Each names its engine pinned by image digest and the exact serve command, in the full record.

Elsewhere in the registry

Registry record as of 2026-09-26. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.