Skip to course content
Free LLMOps course

LLMOps for Reliable AI Applications

Module 06

Agent Evaluation: Tool Trajectory, Approvals, and Stopping

By the end of this module you can evaluate an agent on its trajectory - which tools it called, whether it stopped, whether it asked for approval - not only on its final answer.

Units

  1. Unit 06.00: Judging the path, not just the answer
  2. Unit 06.01: Did it call the right tool, in the right order?
  3. Unit 06.02: Approvals requested and skipped
  4. Unit 06.03: Stopping rules and runs that never end
  5. Unit 06.04: Scoring a trajectory reproducibly

Module work