Skip to course content
Free LLMOps course

LLMOps for Reliable AI Applications

Module 03 Activity

Scenario

There is no eval set. Everything in the rest of this course depends on building one.

What you build

An eval set of at least 30 real cases with recorded sources, a labelling rubric, a holdout, and a process for growing it from failures.

Steps

  1. Collect at least 30 real questions from logs, tickets or colleagues. Record the source of each.
  2. Cover the seven awkward kinds, including messy real input with typos. Report the count per kind and the refusal share.
  3. Write the labelling rubric in terms of properties before labelling anything. Have two people label 10 cases and report agreement.
  4. Split off a holdout you will not tune against, and record the current score on both.
  5. Make three changes, tracking dev and holdout after each. Report the gap.
  6. Write the process that turns an incident into a permanent case, and apply it to one past failure.

Evidence to hand in

Review checklist