models · stepfun-ai
image-text-to-text model harvested from Hugging Face (trending). Catalogued at Quarantine — unverified, pending attestation.
| Method | hf-published-sha256 |
|---|
| Publisher | stepfun-ai |
|---|---|
| Origin | China |
| Licence | apache-2.0 |
| Parameters | 201,365,316,160 (201.4B) |
| Context window | 262,144 tokens |
| Modality | text+image+video->text |
| Native precision | BF16 |
| Pulls recorded here | 50,207 |
| Downloads reported upstream | 19,350 |
| GPQA Diamond | 80.9 |
|---|---|
| IFBench | 67.3 |
| Humanity's Last Exam | 21.4 |
Reported by a third party and recorded here, not re-run by us. Source: artificialanalysis.ai/api/v2.
| Hardware | Engine | Build measured | Decode, 1 stream | Label | Reported by |
|---|---|---|---|---|---|
| Apple M5 Max 128GB | llama.cpp | GGUF IQ3_XXS | 60.3 tok/s @ 8,256 | candidate | exo-postgres |
| Apple M5 Max 128GB | llama.cpp | GGUF iq4xs | 59.6 tok/s @ 8,256 | candidate | exo-postgres |
| Apple M5 Max 128GB | llama.cpp | GGUF IQ4_XS | 59.6 tok/s @ 8,256 | candidate | exo-postgres |
| Apple M3 Ultra 96GB 80-core GPU | llama.cpp | GGUF IQ4_XS | 59.0 tok/s @ 8,256 | candidate | exo-postgres |
| Apple M5 Max 128GB | llama.cpp | GGUF IQ3_XXS | 58.2 tok/s @ 8,256 | candidate | exo-postgres |
| Apple M5 Max 128GB | llama.cpp | Q4_K_S | 58.2 tok/s @ 8,256 | candidate | exo-postgres |
| Apple M3 Ultra 96GB 80-core GPU | llama.cpp | GGUF IQ3_XXS | 57.7 tok/s @ 8,256 | candidate | exo-postgres |
| Apple M3 Ultra 96GB 80-core GPU | llama.cpp | Q4_K_S | 57.7 tok/s @ 8,256 | candidate | exo-postgres |
| Apple M5 Max 128GB | llama.cpp | GGUF UD-IQ2_XXS | 56.5 tok/s @ 8,256 | candidate | exo-postgres |
| Apple M5 Max 128GB | llama.cpp | GGUF Q3_K_M | 56.3 tok/s @ 8,256 | candidate | exo-postgres |
From local-ai-registry (MIT), commit 124959d: 44 runs across 8 machines, 0 validated — model revision and runtime pinned, launch accepted — and 44 candidate, their label for useful evidence without a reproducible-launch promise. 28 more not listed.
| Hardware | Build measured | Decode @ 8K | Memory @ 8K | Decode mode |
|---|---|---|---|---|
| NVIDIA RTX PRO 6000 Blackwell 96GB | IQ3_XXS | 158 tok/s | 47.5 GB | ordinary |
| NVIDIA RTX PRO 6000 Blackwell 96GB | UD-IQ2_XXS | 144 tok/s | 75.6 GB | ordinary |
| NVIDIA RTX PRO 6000 Blackwell 96GB | UD-IQ2_M | 142 tok/s | 75.7 GB | ordinary |
| MacBook Pro M5 Max 128GB · 40-core GPU | IQ3_XXS | 66.5 tok/s | 48.2 GB | ordinary |
| MacBook Pro M5 Max 128GB · 40-core GPU | Q3_K_M | 62.5 tok/s | 106.0 GB | ordinary |
| MacBook Pro M5 Max 128GB · 40-core GPU | UD-IQ2_XXS | 61.9 tok/s | 74.6 GB | ordinary |
| MacBook Pro M5 Max 128GB · 40-core GPU | UD-IQ2_M | 60.1 tok/s | 74.7 GB | ordinary |
| MacBook Pro M5 Max 128GB · 40-core GPU | IQ4_XS | 59.8 tok/s | 117.2 GB | ordinary |
Measured by local.ai, not by us: 33 runs, 25 more not listed. Each names its engine pinned by image digest and the exact serve command, in the full record.
Registry record as of 2026-09-01. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.