models · zai-org

GLM-5.2

text-generation model harvested from Hugging Face (trending). Catalogued at Quarantine — unverified, pending attestation.

Compare Verify

What this registry has checked

Quarantine. We hold the publisher’s own published hash for these weights and have recorded it against this entry. We have not downloaded and hashed the weights ourselves, so the hash is the publisher’s claim, not our measurement.
Methodhf-published-sha256

What is recorded

Publisherzai-org
OriginChina
Licencemit
Parameters753,329,940,480 (753.3B)
Context window1,048,576 tokens
Modalitytext->text
Architectureglm_moe_dsa
Layers78
Hidden size6144
AttentionMLA (64 heads, 64 KV)
Native precisionBF16
Pulls recorded here674
Downloads reported upstream982,017

Reported scores

GPQA Diamond89.5
IFBench73.3
Humanity's Last Exam41.1

Reported by a third party and recorded here, not re-run by us. Source: artificialanalysis.ai/api/v2.

Measured on real hardware

None of these runs are ours. Each figure is the reporter’s, linked to their own evidence, under their own trust label. The number shown is single-stream decode — one user, concurrency 1 — because aggregate throughput across many users is a different quantity and reads as far faster than anyone will see.

Reported to local-ai-registry

HardwareEngineBuild measuredDecode, 1 streamLabelReported by
RTX PRO 6000 Blackwell ×4 devicesvllmEXL3 Trellis TR3 3.0 bpw188 tok/s @ 8,192validated0xsero
Measured outside 2K–32K context — decode speed depends heavily on context, so these are not comparable with the rows above
RTX PRO 6000 Blackwell ×4 devicesvllmModelOpt safetensors NVFP4 non-uniform REAP77.0 tok/scandidate0xsero
RTX PRO 6000 Blackwell ×4 devicesvllmModelOpt safetensors NVFP4 REAP60.0 tok/scandidate0xsero
RTX PRO 6000 Blackwell ×4 devicesvllmsafetensors MXFP8/NVFP4/NF3 hybrid58.2 tok/s @ 180,007candidate0xsero
RTX PRO 6000 Blackwell ×4 devicesllama.cppUD-Q2_K_MXFP447.9 tok/s @ 524,288candidatelocalmaxxing
RTX PRO 6000 Blackwell ×4 devicesllama.cppUD-IQ4_XS39.8 tok/s @ 65,536candidatelocalmaxxing
NVIDIA DGX Spark GB10vllmInt4-Int8Mix38.6 tok/s @ 200,000candidatelocalmaxxing
GeForce RTX 3090 ×10 devicesllama.cppQ3_K_M22.1 tok/s @ 262,144candidatelocalmaxxing
NVIDIA DGX Spark GB10 ×3 devicesvllmhybrid ModelOpt+AQLM NVFP4 hot experts + AQLM 2-bit cold experts21.0 tok/scandidatemialabs

From local-ai-registry (MIT), commit 124959d: 79 runs across 6 machines, 1 validated — model revision and runtime pinned, launch accepted — and 78 candidate, their label for useful evidence without a reproducible-launch promise.

Elsewhere in the registry

Registry record as of 2026-09-01. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.