PurposeAligns fine-tuned models with preference data via DPO, ORPO, KTO, or SimPO. Provides a data-shape method table, iterative on-policy DPO pattern, and μ−2σ pair construction. Outputs a validated method choice and TRL config for direct training use, improving model alignment efficiently.

Domains

Others

Forms

Workflow
Required Tools
None
Languages
python
Package Manager
pipuv
Skill Composition
  • SKILL.md50.0%
  • references50.0%