tools · linkedin
Triton kernels for LLM training that fuse common sequences — RMSNorm, RoPE, SwiGLU, and a chunked fused linear cross-entropy — to cut activation memory and kernel-launch overhead. The fused cross-entropy is the notable one: it avoids materialising the full logits tensor, which is the dominant activation cost at large vocabularies.
| Publisher | |
|---|---|
| Origin | United States |
| Licence | BSD-2-Clause |
Registry record as of 2026-08-14. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.