Module 03 Summary
What this module established
Build from real questions and record the source of each - a case that came from an incident is a failure that became a permanent test. Check the refusal share before trusting any score: at zero, the set cannot detect an over-eager assistant however many cases it has.
Carry forward
- Deliberately include the cases that lower the score, including messy real input. A set that always passes is not a measurement.
- Write the rubric in terms of properties before labelling, and keep 'incomplete' distinct from 'wrong'.
- Keep a holdout and look at it rarely. A widening dev-holdout gap means your improvements are fitting the cases rather than the task.
Before moving on
Move on when you have 30+ real cases with sources, a rubric two labellers agree on, and a holdout you have not tuned against.
