models · nvidia
Catalogued from llm-stats Open LLM Leaderboard at Quarantine — unverified, pending attestation.
| Method | hf-published-sha256 |
|---|
| Publisher | nvidia |
|---|---|
| Origin | United States |
| Licence | unknown |
| Publisher’s own licence statement | other |
| Parameters | 67,228,556,288 (67.2B) |
| Architecture | nemotron_h |
| Layers | 88 |
| Hidden size | 4096 |
| Attention | GQA (32 heads, 2 KV) |
| Native precision | U8 |
| Pulls recorded here | 6 |
| Downloads reported upstream | 2,790,020 |
| Hardware | Build measured | Decode @ 8K | Memory @ 8K | Decode mode |
|---|---|---|---|---|
| NVIDIA RTX PRO 6000 Blackwell 96GB | NVFP4 | 97.0 tok/s | 95.3 GB | ordinary |
| NVIDIA DGX Spark 128GB | NVFP4 | 15.6 tok/s | 97.4 GB | ordinary |
Measured by local.ai, not by us: 2 runs. Each names its engine pinned by image digest and the exact serve command, in the full record.
Registry record as of 2026-09-01. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.