Module 03 Activity
Scenario
There is no eval set. Everything in the rest of this course depends on building one.
What you build
An eval set of at least 30 real cases with recorded sources, a labelling rubric, a holdout, and a process for growing it from failures.
Steps
- Collect at least 30 real questions from logs, tickets or colleagues. Record the source of each.
- Cover the seven awkward kinds, including messy real input with typos. Report the count per kind and the refusal share.
- Write the labelling rubric in terms of properties before labelling anything. Have two people label 10 cases and report agreement.
- Split off a holdout you will not tune against, and record the current score on both.
- Make three changes, tracking dev and holdout after each. Report the gap.
- Write the process that turns an incident into a permanent case, and apply it to one past failure.
Evidence to hand in
- The 30+ cases with kind and source.
- The kind counts and refusal share.
- The rubric and the two-labeller agreement figure.
- Dev and holdout scores across three changes, with the gap.
- The incident-to-case process and one case produced by it.
Review checklist
- Cases came from real usage, not imagination.
- The refusal share is at least a quarter.
- The rubric was written before labelling.
- The holdout was looked at rarely and on purpose.
- At least one case came from a real past failure.
