models · cohere
North Mini Code — catalogued from third-party evaluation data; no weights held by Sovereign Frontier.
| Method | hf-published-sha256 |
|---|
| Publisher | cohere |
|---|---|
| Origin | Canada |
| Licence | Apache-2.0 |
| Parameters | 30,484,303,872 (30.5B) |
| Context window | 256,000 tokens |
| Modality | text->text |
| Architecture | cohere2_moe |
| Layers | 49 |
| Hidden size | 2048 |
| Attention | GQA (32 heads, 4 KV) |
| Native precision | BF16 |
| Pulls recorded here | 8 |
| Downloads reported upstream | 11,855 |
| GPQA Diamond | 75.7 |
|---|---|
| IFBench | 57.6 |
| Humanity's Last Exam | 11.1 |
Reported by a third party and recorded here, not re-run by us. Source: artificialanalysis.ai/api/v2.
| Hardware | Engine | Build measured | Decode, 1 stream | Label | Reported by |
|---|---|---|---|---|---|
| GeForce RTX 5090 | llama.cpp | UD-Q6_K_XL | 258 tok/s @ 8,192 | candidate | localmaxxing |
| Measured outside 2K–32K context — decode speed depends heavily on context, so these are not comparable with the rows above | |||||
| GeForce RTX 3090 | llama.cpp | Q4_K_M | 174 tok/s @ 512 | candidate | localmaxxing |
| GeForce RTX 3060 ×2 devices | llama.cpp | Q4_K_M | 102 tok/s @ 512 | candidate | localmaxxing |
From local-ai-registry (MIT), commit 124959d: 3 runs across 3 machines, 0 validated — model revision and runtime pinned, launch accepted — and 3 candidate, their label for useful evidence without a reproducible-launch promise.
Registry record as of 2026-09-26. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.