tools · flashinfer-ai
Attention kernel library built for inference serving rather than training: paged and ragged KV-cache layouts, batched prefill and decode kernels, and CUDA-graph-friendly execution. Adopted as a kernel backend by several serving stacks, so it sits underneath the engines rather than beside them.
| Publisher | flashinfer-ai |
|---|---|
| Origin | United States |
| Licence | Apache-2.0 |
Registry record as of 2026-08-14. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.