Skip to course content
Free LLMOps course

LLMOps for Reliable AI Applications

Module 03

Eval Datasets and Test-Case Design

By the end of this module you can build an evaluation dataset that reflects real usage, including the awkward cases people actually send rather than the ones you hoped for.

Units

  1. Unit 03.00: Building a set from real usage
  2. Unit 03.01: The awkward cases people actually send
  3. Unit 03.02: Labelling without arguing about it later
  4. Unit 03.03: Keeping a holdout you do not tune against
  5. Unit 03.04: Growing the set from production failures

Module work