PurposeConverts graded evaluation traces and production logs into SFT examples and DPO preference pairs for LLM fine-tuning. Uses reward ranking, expert-correction routing, step-level masking, and same-task pair building with μ−2σ rejection. Outputs hygiene-checked JSONL that feeds directly into dataset-curation, closing the eval-to-training flywheel.

Domains

Others

Forms

Workflow
Required Tools
None
Languages
python
Package Manager
uv
Skill Composition
  • SKILL.md50.0%
  • references50.0%