tools · bytedance
Reinforcement-learning library for post-training language models, built around a hybrid controller that separates the RL control flow from the distributed compute, so rollout, reward and update stages can each use a different parallelism strategy and engine. Used for RLHF and verifiable-reward pipelines.
| Publisher | bytedance |
|---|---|
| Origin | China |
| Licence | Apache-2.0 |
Registry record as of 2026-08-14. Identity, licence and parameter count are properties of a release and do not move; tier can, and the verify link above re-checks it against the live registry rather than this page.