models · zai-org
text-generation model harvested from Hugging Face (trending). Catalogued at Quarantine — unverified, pending attestation.
| Method | hf-published-sha256 |
|---|
| Publisher | zai-org |
|---|---|
| Origin | China |
| Licence | mit |
| Parameters | 753,329,940,480 (753.3B) |
| Context window | 1,048,576 tokens |
| Modality | text->text |
| Architecture | glm_moe_dsa |
| Layers | 78 |
| Hidden size | 6144 |
| Attention | MLA (64 heads, 64 KV) |
| Native precision | BF16 |
| Pulls recorded here | 674 |
| Downloads reported upstream | 982,017 |
| GPQA Diamond | 89.5 |
|---|---|
| IFBench | 73.3 |
| Humanity's Last Exam | 41.1 |
Reported by a third party and recorded here, not re-run by us. Source: artificialanalysis.ai/api/v2.
| Hardware | Engine | Build measured | Decode, 1 stream | Label | Reported by |
|---|---|---|---|---|---|
| RTX PRO 6000 Blackwell ×4 devices | vllm | EXL3 Trellis TR3 3.0 bpw | 188 tok/s @ 8,192 | validated | 0xsero |
| Measured outside 2K–32K context — decode speed depends heavily on context, so these are not comparable with the rows above | |||||
| RTX PRO 6000 Blackwell ×4 devices | vllm | ModelOpt safetensors NVFP4 non-uniform REAP | 77.0 tok/s | candidate | 0xsero |
| RTX PRO 6000 Blackwell ×4 devices | vllm | ModelOpt safetensors NVFP4 REAP | 60.0 tok/s | candidate | 0xsero |
| RTX PRO 6000 Blackwell ×4 devices | vllm | safetensors MXFP8/NVFP4/NF3 hybrid | 58.2 tok/s @ 180,007 | candidate | 0xsero |
| RTX PRO 6000 Blackwell ×4 devices | llama.cpp | UD-Q2_K_MXFP4 | 47.9 tok/s @ 524,288 | candidate | localmaxxing |
| RTX PRO 6000 Blackwell ×4 devices | llama.cpp | UD-IQ4_XS | 39.8 tok/s @ 65,536 | candidate | localmaxxing |
| NVIDIA DGX Spark GB10 | vllm | Int4-Int8Mix | 38.6 tok/s @ 200,000 | candidate | localmaxxing |
| GeForce RTX 3090 ×10 devices | llama.cpp | Q3_K_M | 22.1 tok/s @ 262,144 | candidate | localmaxxing |
| NVIDIA DGX Spark GB10 ×3 devices | vllm | hybrid ModelOpt+AQLM NVFP4 hot experts + AQLM 2-bit cold experts | 21.0 tok/s | candidate | mialabs |
From local-ai-registry (MIT), commit 124959d: 79 runs across 6 machines, 1 validated — model revision and runtime pinned, launch accepted — and 78 candidate, their label for useful evidence without a reproducible-launch promise.
Registry record as of 2026-09-01. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.