models · zhipu-ai

GLM-4.7-Flash

Zhipu AI GLM-4.7-Flash. 30B compact variant. Arena 759. GPQA Diamond 75.2%, AIME 91.6%. Low-latency deployment of GLM-4 capabilities.

Compare Verify

What this registry has checked

Quarantine. We hold the publisher’s own published hash for these weights and have recorded it against this entry. We have not downloaded and hashed the weights ourselves, so the hash is the publisher’s claim, not our measurement.
Weight file SHA-2564443c1888927bd19c3728a10043a1cd9a2369430efef192ffa70fc22f7be12c2
Methodhf-published-sha256

What is recorded

Publisherzhipu-ai
OriginRecorded as a covered nation under 10 U.S.C. § 4872(f) — the PRC, Russia, Iran or North Korea. The registry did not record which.
LicenceMIT
Parameters31,221,488,576 (31.2B)
Context window128K
Modalitytext->text
Architectureglm4_moe_lite
Layers47
Hidden size2048
AttentionMLA (20 heads, 20 KV)
Native precisionBF16
Pulls recorded here290,017
Downloads reported upstream1,858,331

Reported scores

GPQA Diamond58.1
IFBench60.8
Humanity's Last Exam7.6

Reported by a third party and recorded here, not re-run by us. Source: artificialanalysis.ai/api/v2.

Measured on real hardware

None of these runs are ours. Each figure is the reporter’s, linked to their own evidence, under their own trust label. The number shown is single-stream decode — one user, concurrency 1 — because aggregate throughput across many users is a different quantity and reads as far faster than anyone will see.

Reported to local-ai-registry

HardwareEngineBuild measuredDecode, 1 streamLabelReported by
GeForce RTX 5090llama.cppUD-Q6_K_XL176 tok/s @ 8,192candidatelocalmaxxing
Radeon AI PRO R9700llama.cppQ2_K_XL82.0 tok/s @ 2,048candidatelocalmaxxing
Radeon AI PRO R9700llama.cppQ4_K_XL75.8 tok/s @ 2,048candidatelocalmaxxing
Ryzen AI Max+ 395llama.cppQ4_K_XL53.0 tok/s @ 4,096candidatelocalmaxxing
Intel Arc Pro B70llama.cppGGUF UD-Q4_K_XL40.8 tok/s @ 4,096candidatelocalmaxxing
Measured outside 2K–32K context — decode speed depends heavily on context, so these are not comparable with the rows above
GeForce RTX 5090llama.cppQ5_K_M219 tok/s @ 131,072candidatelocalmaxxing
Radeon AI PRO R9700 ×3 devicesllama.cppMXFP4_MOE97.4 tok/s @ 783candidatelocalmaxxing

From local-ai-registry (MIT), commit 124959d: 13 runs across 6 machines, 0 validated — model revision and runtime pinned, launch accepted — and 13 candidate, their label for useful evidence without a reproducible-launch promise.

Elsewhere in the registry

Registry record as of 2026-09-26. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.