PurposePrepares, formats, and validates datasets for supervised fine-tuning and preference training. Methods: format selection by target method, chat template application with loss masking, sequence packing, synthetic data mixing with ≥25% real floor, and dataset card authoring. Produces a validated JSONL dataset and Phase 2 card that gates the training run.

Domains

Others

Forms

Workflow
Required Tools
None
Languages
python
Package Manager
N/A
Skill Composition
  • SKILL.md33.3%
  • references66.7%