models · poolside
Catalogued from local.ai local-inference benchmark index at Quarantine — unverified, pending attestation.
| Method | hf-published-sha256 |
|---|
| Publisher | poolside |
|---|---|
| Origin | Not recorded. The harvest wrote its own source name into this field, which tells us nothing about jurisdiction. |
| Licence | OpenMDW-1.1 |
| Parameters | 117,561,977,600 (117.6B) |
| Context window | 262,144 tokens |
| Modality | text->text |
| Architecture | laguna |
| Layers | 48 |
| Hidden size | 3072 |
| Attention | GQA (48 heads, 8 KV) |
| Native precision | BF16 |
| Downloads reported upstream | 64,991 |
| Hardware | Engine | Build measured | Decode, 1 stream | Label | Reported by |
|---|---|---|---|---|---|
| Radeon AI PRO R9700 ×3 devices | llama.cpp | Q4_0_ROCMFP4_COHERENT | 52.3 tok/s @ 2,048 | candidate | localmaxxing |
| Radeon AI PRO R9700 ×3 devices | llama.cpp | Q4_0_ROCMFP4_STRIX_LEAN | 52.3 tok/s @ 2,048 | candidate | localmaxxing |
| Radeon AI PRO R9700 ×3 devices | llama.cpp | Q4_0_ROCMFP4_FAST | 52.2 tok/s @ 2,048 | candidate | localmaxxing |
| Measured outside 2K–32K context — decode speed depends heavily on context, so these are not comparable with the rows above | |||||
| Radeon AI PRO R9700 ×3 devices | llama.cpp | UD-Q4_K_XL | 51.6 tok/s @ 131,072 | candidate | localmaxxing |
| Ryzen AI Max+ 395 | llama.cpp | Q4_0_ROCMFP4_STRIX_LEAN | 41.7 tok/s @ 262,144 | candidate | localmaxxing |
| Ryzen AI Max+ 395 | llama.cpp | Q4_0_ROCMFP4_FAST | 40.7 tok/s @ 820 | candidate | localmaxxing |
| Ryzen AI Max+ 395 | llama.cpp | UD-Q4_K_XL | 29.9 tok/s @ 820 | candidate | localmaxxing |
From local-ai-registry (MIT), commit 124959d: 8 runs across 3 machines, 0 validated — model revision and runtime pinned, launch accepted — and 8 candidate, their label for useful evidence without a reproducible-launch promise.
| Hardware | Build measured | Decode @ 8K | Memory @ 8K | Decode mode |
|---|---|---|---|---|
| NVIDIA RTX PRO 6000 Blackwell 96GB | Q4_K_M | 117 tok/s | 80.1 GB | ordinary |
| MacBook Pro M5 Max 128GB · 40-core GPU | Q4_K_M | 48.3 tok/s | 87.0 GB | ordinary |
| Mac Studio M3 Ultra 96GB · 80-core GPU | Q4_K_M | 47.4 tok/s | 88.7 GB | ordinary |
| Mac Studio M3 Ultra 96GB · 60-core GPU | Q4_K_M | 47.1 tok/s | 78.1 GB | ordinary |
| NVIDIA DGX Spark 128GB | Q4_K_M | 20.1 tok/s | 86.6 GB | ordinary |
| NVIDIA DGX Spark 128GB | NVFP4 | 18.5 tok/s | 73.3 GB | ordinary |
Measured by local.ai, not by us: 6 runs. Each names its engine pinned by image digest and the exact serve command, in the full record.
Registry record as of 2026-09-26. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.