tools · huggingface

TRL

Transformer Reinforcement Learning — PPO, DPO, SFT Trainer for RLHF pipelines.

What this registry has checked

Quarantine. We hold the publisher’s own published hash for these weights and have recorded it against this entry. We have not downloaded and hashed the weights ourselves, so the hash is the publisher’s claim, not our measurement.
Weight file SHA-25629825e1ad16188eb350f5761523cc903d6c79a330d426a05a124ea1bbdcd0f92

What is recorded

Publisherhuggingface
OriginInternational
LicenceApache-2.0
Pulls recorded here96,001

Elsewhere in the registry

Registry record as of 2026-06-06. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.