PurposeTrain reasoning and verifiable-task behavior using GRPO and RLVR. Provides a TRL GRPOTrainer recipe, composite reward design, a mandatory reward-inspection gate, and variant selection. Produces validated GRPO configs that reduce reward-hacking and sharpen models on algorithmically checkable tasks.

Domains

Coding

Forms

Workflow
Required Tools
None
Languages
python
Package Manager
pipuv
Skill Composition
  • SKILL.md33.3%
  • references66.7%