PurposeImplements comprehensive evaluation strategies for LLM apps via automated metrics, human feedback, and LLM-as-Judge. Provides Python suite to score outputs on test cases. Helps detect regressions, compare models, and build production confidence.

Domains

CodingOthers

Forms

Workflow
Required Tools
None
Languages
python
Package Manager
N/A
Skill Composition
  • SKILL.md50.0%
  • references50.0%