SOTA · NanoGPT speedrun · val loss 3.28, 124M params, 8×H200
Every submission rerun under signed conditions on the sealed cluster SFC-7.
loading...
loading...
click to expand · verify re-checks sigs + inclusion
Optimizers · live
Optimizer implementations and learning-rate schedules. Ranked by wall-clock time to reach a fixed target loss on identical hardware — the honest measure, since steps-to-converge hides per-step cost. Served live from .
SOTA · NanoGPT speedrun · val loss 3.28, 124M params, 8×H200
Every submission rerun under signed conditions on the sealed cluster SFC-7.