models · nvidia
Cheapest input model on the leaderboard ($0.06/1M tok). NVIDIA 30B sparse MoE optimized for cost-efficiency.
| Weight file SHA-256 | 5bb70b90ff8115e32cebce8e66da6e57692e25cb6df4a722e62750f0410b5939 |
|---|---|
| Method | hf-published-sha256 |
| Publisher | nvidia |
|---|---|
| Origin | United States |
| Licence | Apache-2.0 |
| Publisher’s own licence statement | other |
| Parameters | 31,577,940,288 (31.6B) |
| Context window | 262,144 tokens |
| Modality | text->text |
| Architecture | nemotron_h |
| Layers | 52 |
| Hidden size | 2688 |
| Attention | GQA (32 heads, 2 KV) |
| Native precision | BF16 |
| Pulls recorded here | 41,006 |
| Downloads reported upstream | 745,812 |
| Hardware | Engine | Build measured | Decode, 1 stream | Label | Reported by |
|---|---|---|---|---|---|
| GeForce RTX 5090 | llama.cpp | UD-Q4_K_XL | 313 tok/s @ 8,192 | candidate | localmaxxing |
| Measured outside 2K–32K context — decode speed depends heavily on context, so these are not comparable with the rows above | |||||
| GeForce RTX 5090 | ollama | Q4_K_M | 286 tok/s @ 800 | candidate | localmaxxing |
| GeForce RTX 3060 ×2 devices | llama.cpp | Q4_K_S | 86.2 tok/s @ 262,144 | candidate | localmaxxing |
From local-ai-registry (MIT), commit 124959d: 6 runs across 3 machines, 0 validated — model revision and runtime pinned, launch accepted — and 6 candidate, their label for useful evidence without a reproducible-launch promise.
Registry record as of 2026-09-01. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.