Domains
Forms
Domains
Forms
Mechanism
Breaks down how the skill works and what it produces, so you can quickly judge whether it fits your scenario and value.
When starting any fine-tuning effort, unsure whether RAG/prompting suffices, or choosing between preference-optimization and reinforcement methods.
A clear routing decision: whether to fine-tune, which method (SFT/DPO/KTO/GRPO/RLVR/CPT), and base-model size class — preventing wasted training runs and wrong method choices.
Use Cases
A team wants the assistant to follow internal support macros exactly. They are unsure whether to fine-tune or just prompt. Using this skill, they confirm behavior is stable and demonstrable from transcripts, route to SFT via lora-qlora-recipes, and skip costly RAG/prompt iteration. Value: correct method chosen fast, training effort focused where it pays off.
A researcher has auto-gradable math problems and wants reasoning improvement. Unsure between DPO and GRPO. This skill shows verifiable success signals route to GRPO+RLVR, and only after the model succeeds sometimes. They hand off to grpo-rlvr-training with right expectations. Value: avoids wrong DPO run and teaches when RL is viable.
Skill Relationships
Dependency relationships read as "the upper tier points to the lower tier." The current Skill sits in the middle tier — above are Skills that depend on it, below are Skills it depends on.
Tier 1 · These Skills Use Me
Tier 2 · Current Skill
Tier 3 · I Use These Skills
Browse this skill's relationships within its skillset. Click a node to switch the side panel; use the search box to jump to any skill.
Skill File