eval-harness-first
Maintained by wshobson
Build the evaluation harness that gates every fine-tuning run — golden sets, per-failure-mode graders, judge calibration, and base-model baselines. Use when starting a fine-tuning effort, when converting traces into an eval set, or when calibrating a judge against human labels.
- Current version
- Unknown
- License
- Unknown
- Network access
- Unknown / not assessed
- Review status
- Not verified
Problem it solves
This catalog entry helps users find and evaluate eval-harness-first for the task described by its available catalog summary. Confirm the exact scope in the linked original source when one is available.
When to use it
Consider eval-harness-first when its available catalog summary matches the task at hand. When available, review the linked original source before use for precise instructions, requirements, and limitations.
Installation and updates
These commands are displayed for copying only and are never executed on RefHub servers. Review the linked upstream source before running them.
npx skills add wshobson/agents --skill eval-harness-first -y
Agent compatibility
No compatibility test has been recorded
Do not assume agent compatibility until documented test evidence is available.