models · deepseek-ai

DeepSeek-V4-Flash-0731

text-generation model harvested from Hugging Face (trending). Catalogued at Quarantine — unverified, pending attestation.

Compare Verify

What this registry has checked

Quarantine. We hold the publisher’s own published hash for these weights and have recorded it against this entry. We have not downloaded and hashed the weights ourselves, so the hash is the publisher’s claim, not our measurement.
Methodhf-published-sha256

What is recorded

Publisherdeepseek-ai
OriginRecorded as a covered nation under 10 U.S.C. § 4872(f) — the PRC, Russia, Iran or North Korea. The registry did not record which.
Licencemit
Parameters304,180,418,494 (304.2B)
Context window1,310,720 tokens
Modalitytext->text
Architecturedeepseek_v4
Layers43
Hidden size4096
AttentionGQA (64 heads, 1 KV)
Native precisionI8
Pulls recorded here15,372
Downloads reported upstream4,469,509

Reported scores

GPQA Diamond90.8
Humanity's Last Exam38.6

Reported by a third party and recorded here, not re-run by us. Source: artificialanalysis.ai/api/v2.

Measured on real hardware

None of these runs are ours. Each figure is the reporter’s, linked to their own evidence, under their own trust label. The number shown is single-stream decode — one user, concurrency 1 — because aggregate throughput across many users is a different quantity and reads as far faster than anyone will see.

Reported to local-ai-registry

HardwareEngineBuild measuredDecode, 1 streamLabelReported by
Radeon AI PRO R9700llama.cppMXFP452.0 tok/s @ 32,768candidatelocalmaxxing
Radeon AI PRO R9700 ×3 devicesllama.cppAD-IQ1_M48.5 tok/s @ 32,768candidatelocalmaxxing
NVIDIA DGX Spark GB10 ×2 devicesvllmModelOpt NVFP443.7 tok/s @ 8,192candidatemialabs
GeForce RTX 3090llama.cppIQ2_XSS20.7 tok/s @ 2,048candidatelocalmaxxing
GeForce RTX 5090llama.cppUD-IQ1_S18.0 tok/s @ 2,048candidatelocalmaxxing
GeForce RTX 5090llama.cppIQ1_S17.5 tok/s @ 2,048candidatelocalmaxxing
GeForce RTX 5090llama.cppQ8_K_XL10.4 tok/s @ 2,048candidatelocalmaxxing
Measured outside 2K–32K context — decode speed depends heavily on context, so these are not comparable with the rows above
RTX PRO 6000 Blackwell ×2 devicesvllmsafetensors FP8_E4M3276 tok/s @ 262,144candidatelocalmaxxing
RTX PRO 6000 Blackwell ×2 devicessglangsafetensors FP857.8 tok/s @ 50candidateppickle1989
GeForce RTX 3090 ×7 devicesllama.cppUD-IQ4_XS50.8 tok/s @ 262,144candidatelocalmaxxing

From local-ai-registry (MIT), commit 124959d: 73 runs across 6 machines, 0 validated — model revision and runtime pinned, launch accepted — and 73 candidate, their label for useful evidence without a reproducible-launch promise. 5 more not listed.

Measured by local.ai

HardwareBuild measuredDecode @ 8KMemory @ 8KDecode mode
MacBook Pro M5 Max 128GB · 40-core GPUUD-Q2_K_XL29.9 tok/s99.1 GBordinary
NVIDIA DGX Spark 128GBUD-IQ2_M19.9 tok/s93.1 GBordinary
NVIDIA DGX Spark 128GBUD-IQ3_XXS18.8 tok/s105.5 GBordinary
NVIDIA DGX Spark 128GBUD-Q2_K_XL17.5 tok/s98.8 GBordinary
Mac Studio M3 Ultra 96GB · 80-core GPUUD-Q2_K_XL7.1 tok/s99.1 GBordinary

Measured by local.ai, not by us: 5 runs. Each names its engine pinned by image digest and the exact serve command, in the full record.

Elsewhere in the registry

Registry record as of 2026-09-01. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.