PurposeFine-tunes vision-language models (VLMs) with supervised learning on image+text data. Provides consensus recipe (frozen vision tower + LLM LoRA), escalation rules for visual domain shift, and two silent-killer checks. Produces a validated adapter config that prevents non-learning runs and feeds directly into training script generation.

Domains

Coding

Forms

Style GuideWorkflow
Required Tools
None
Languages
python
Package Manager
N/A
Skill Composition
  • SKILL.md50.0%
  • references50.0%