Module 06 Activity
Scenario
Your assistant answers. The question now is what it does when it should not - the behaviour no set of real questions will test on its own.
What you produce
A refusal specification and an evaluation set in which a substantial share of cases expect a refusal.
Steps
- Fix the refusal wording exactly, and specify that it carries no citation and is recorded as a refusal rather than an answer.
- Add five unanswerable questions: three the corpus genuinely lacks, and two it covers but phrased in a way you expect retrieval to miss.
- For every case, write down the decision you expect - answer or refuse - before looking at anything.
- Separate the two ways a refusal decision can be wrong, and say which is more expensive for your users and why.
- Write one bounded partial answer for a compound question: the half you have, and a sentence naming the half you do not.
- Say what proportion of your set expects a refusal, and whether that is enough to detect an over-eager assistant.
Evidence to hand in
- The refusal wording and what accompanies it.
- The evaluation set with an expected decision on every case.
- The two failure directions with your cost judgement.
- One bounded partial answer.
- The refusal share.
Review checklist
- The refusal wording is exact and could be matched by a check.
- The refusal share is at least a quarter of the set.
- The two cases the corpus covers are identified as retrieval problems, not content gaps.
- The two failure directions are treated separately, not summed.
- The partial answer names what was not covered.
