optimizers · google
Layer-wise Adaptive Moments for large-batch training (You et al., 2019). Enabled BERT pretraining in 76 minutes at batch size 32k.
| Weight file SHA-256 | b1b57872d4d3d9f967dae7e540ab5c3eeae2e235a826b0eb160352bbdf753942 |
|---|
| Publisher | |
|---|---|
| Origin | United States |
| Licence | Apache-2.0 |
| Pulls recorded here | 44,302 |
Registry record as of 2026-06-06. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.