models · inclusionai
text-generation model harvested from Hugging Face (trending). Catalogued at Quarantine — unverified, pending attestation.
| Method | hf-published-sha256 |
|---|
| Publisher | inclusionai |
|---|---|
| Origin | China |
| Licence | mit |
| Parameters | 127,486,405,600 (127.5B) |
| Context window | 262,144 tokens |
| Modality | text->text |
| Architecture | bailing_hybrid |
| Layers | 42 |
| Hidden size | 2560 |
| Attention | MLA (32 heads, 32 KV) |
| Native precision | BF16 |
| Pulls recorded here | 1,196 |
| Downloads reported upstream | 17,926 |
| GPQA Diamond | 85.5 |
|---|---|
| Humanity's Last Exam | 23.7 |
Reported by a third party and recorded here, not re-run by us. Source: artificialanalysis.ai/api/v2.
| Hardware | Engine | Build measured | Decode, 1 stream | Label | Reported by |
|---|---|---|---|---|---|
| Radeon AI PRO R9700 ×3 devices | llama.cpp | Q4_0_ROCMFP4_STRIX | 39.3 tok/s @ 4,096 | candidate | localmaxxing |
| Radeon AI PRO R9700 ×3 devices | llama.cpp | Q4_0_ROCMFP4_STRIX_LEAN | 39.3 tok/s @ 4,096 | candidate | localmaxxing |
| Ryzen AI Max+ 395 | llama.cpp | Q4_0_ROCMFP4_STRIX | 39.2 tok/s @ 4,096 | candidate | localmaxxing |
| Radeon AI PRO R9700 ×3 devices | llama.cpp | Q4_K_M | 33.6 tok/s @ 4,096 | candidate | localmaxxing |
From local-ai-registry (MIT), commit 124959d: 4 runs across 2 machines, 0 validated — model revision and runtime pinned, launch accepted — and 4 candidate, their label for useful evidence without a reproducible-launch promise.
Registry record as of 2026-09-26. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.