Module 08 Activity
Scenario
You have been changing the chunking, the number of passages, the rewriting and the wording for six modules. Without a recorded baseline you cannot say whether any of it helped.
What you produce
An evaluation plan: a set built from kinds, three things scored separately, and a per-case comparison method.
Steps
- Grow the set to at least twenty cases covering all seven kinds. Record where each came from.
- For each case, specify three separate judgements: was the right passage retrieved, was the answer supported by its citation, and was the answer correct.
- Score your fifteen from Module 3 by hand on all three and record the results per case.
- Check your Module 3 prediction: which questions did you expect to fail, and which actually did?
- Make one deliberate change - the number of passages, the chunk grain, or the wording - and re-score.
- Report the fixed and the broken lists rather than the headline rate, and say whether you would keep the change.
Evidence to hand in
- The evaluation set with kind and source per case.
- The three-way scoring definition.
- Per-case results for both runs.
- Your prediction against what happened.
- The fixed and broken lists, and the keep-or-revert decision.
Review checklist
- Retrieval, citation support and correctness are scored separately.
- Per-case results are recorded for both runs.
- The decision cites the broken list, not the summary rate.
- At least one case the system currently fails is kept in the set.
- The Module 3 prediction is checked honestly, including where it was wrong.
