Evaluation Sets for RAG
By the end of this module you can build an evaluation set for a RAG system, testing retrieval and answer quality separately so you know which half is failing.
Units
- Unit 08.00: Building a question set from real usage
- Unit 08.01: Judging retrieval separately from the answer
- Unit 08.02: Scoring citations, not just correctness
- Unit 08.03: The cases you would rather leave out
- Unit 08.04: Re-running the set after every change
