Score Output With a Simple Rubric
Unit ID: PAW-M08-U03 Estimated active time: 20-25 minutes
Scoring makes review comparable
An unscored review produces a feeling. A scored one produces a record you can compare across outputs, across prompts, and across time.
That comparison is what tells you whether your prompt changes are actually working.
Keep the scale small
Three levels is enough: weak, acceptable, strong.
Five-point scales invite the middle, and the middle carries no information. Three forces a decision.
Score against criteria you wrote first
A rubric written after reading the output will describe that output.
Write the criteria with the prompt โ Module 6 covered this โ and score against them unchanged. If the criteria turn out to be wrong, change them for next time rather than mid-review.
The scoring sheet
SCORING SHEET
Output: (which one) Prompt version: (which) Reviewer: (who)
Criterion Weak Acceptable Strong Note
------------- ---- ---------- ------ -----------------------------
Accuracy x one figure unverified
Completeness x
Usefulness x no next action stated
Honesty x assumptions listed and ranked
Tone x
Lowest criterion: Usefulness
One change for next time: add "end with the next action and who owns it"
The two lines at the bottom are the whole point. A score with no resulting change is administration.
Track the lowest criterion over time
If usefulness is the lowest score three times running, that is a prompt problem, not an output problem.
Fixing the prompt once beats editing the output every time โ and the log from Module 7 is where that fix gets recorded.
Mini practice
Score your last three outputs on the same five criteria.
Find the criterion that scored lowest most often.
Write the one prompt line that would address it, and add it today.
