Evaluating an Agentic Workflow
By the end of this module you can evaluate an agentic workflow on what it did, not just what it produced: which tools it called, in what order, and where it should have stopped.
Units
- Unit 08.00: What you are actually measuring
- Unit 08.01: Building the case set from real requests
- Unit 08.02: Judging the path as well as the result
- Unit 08.03: The cases that make the score look worse
- Unit 08.04: Re-running the set after every change
