models · openai

GPT OSS 120B

OpenAI's 120B open-source model released under Apache-2.0. Strong general reasoning and coding.

Compare Verify

What this registry has checked

Quarantine. We hold the publisher’s own published hash for these weights and have recorded it against this entry. We have not downloaded and hashed the weights ourselves, so the hash is the publisher’s claim, not our measurement.
Weight file SHA-256fee26c4f4e06ca8f9dcb5b135eaeacb732eb10eaae91b5076e2b7fabdf4336e6
Methodhf-published-sha256

What is recorded

Publisheropenai
OriginUnited States
LicenceApache-2.0
Parameters116,829,156,672 (116.8B)
Context window131,072 tokens
Modalitytext->text
Architecturegpt_oss
Layers36
Hidden size2880
AttentionGQA (64 heads, 8 KV)
Native precisionU8
Pulls recorded here156,019
Downloads reported upstream5,378,443

Reported scores

MMLU-Pro80.8
GPQA Diamond78.2
AIME 202593.4
LiveCodeBench87.8
IFBench69
Humanity's Last Exam19.6

Reported by a third party and recorded here, not re-run by us. Source: artificialanalysis.ai/api/v2.

Measured on real hardware

None of these runs are ours. Each figure is the reporter’s, linked to their own evidence, under their own trust label. The number shown is single-stream decode — one user, concurrency 1 — because aggregate throughput across many users is a different quantity and reads as far faster than anyone will see.

Reported to local-ai-registry

HardwareEngineBuild measuredDecode, 1 streamLabelReported by
Intel Arc Pro B70 ×2 devicesllama.cppMXFP459.4 tok/s @ 2,048candidatelocalmaxxing
NVIDIA DGX Spark GB10vllmMXFP434.4 tok/s @ 32,768candidatelocalmaxxing
Measured outside 2K–32K context — decode speed depends heavily on context, so these are not comparable with the rows above
RTX PRO 6000 Blackwellllama.cppF16203 tok/s @ 845candidatelocalmaxxing
Radeon AI PRO R9700 ×3 devicesllama.cppF1672.8 tok/s @ 845candidatelocalmaxxing
Ryzen AI Max+ 395llama.cppF1638.8 tok/s @ 845candidatelocalmaxxing

From local-ai-registry (MIT), commit 124959d: 6 runs across 5 machines, 0 validated — model revision and runtime pinned, launch accepted — and 6 candidate, their label for useful evidence without a reproducible-launch promise.

Measured by local.ai

HardwareBuild measuredDecode @ 8KMemory @ 8KDecode mode
NVIDIA RTX PRO 6000 Blackwell 96GB—137 tok/s90.9 GBnot stated
NVIDIA RTX PRO 6000 Blackwell 96GB—128 tok/s94.1 GBnot stated
NVIDIA RTX PRO 6000 Blackwell 96GB—124 tok/s94.1 GBnot stated
MacBook Pro M5 Max 128GB · 40-core GPUMXFP471.3 tok/s64.8 GBnot stated
Mac Studio M3 Ultra 96GB · 80-core GPUMXFP467.9 tok/s64.8 GBnot stated
MacBook Pro M4 Max 128GB · 40-core GPUMXFP462.1 tok/s64.8 GBnot stated
NVIDIA DGX Spark 128GB—33.7 tok/s102.8 GBnot stated
NVIDIA DGX Spark 128GB—33.0 tok/s88.7 GBnot stated

Measured by local.ai, not by us: 16 runs, 8 more not listed. Each names its engine pinned by image digest and the exact serve command, in the full record.

Elsewhere in the registry

Registry record as of 2026-08-13. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.