tools · huggingface
Transformer Reinforcement Learning — PPO, DPO, SFT Trainer for RLHF pipelines.
| Weight file SHA-256 | 29825e1ad16188eb350f5761523cc903d6c79a330d426a05a124ea1bbdcd0f92 |
|---|
| Publisher | huggingface |
|---|---|
| Origin | International |
| Licence | Apache-2.0 |
| Pulls recorded here | 96,001 |
Registry record as of 2026-06-06. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.