systems · nvidia
Datacentre-scale inference serving framework built around disaggregated serving: prefill and decode run on separate worker pools so each scales and schedules independently, with a KV-cache-aware router and a distributed KV store to move cache between them. Engine-agnostic — it sits above vLLM, TensorRT-LLM and SGLang rather than replacing them.
| Publisher | nvidia |
|---|---|
| Origin | United States |
| Licence | Apache-2.0 |
| Pulls recorded here | 2 |
Registry record as of 2026-08-14. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.