Practice: Evaluate Two Outputs for the Same Task
Unit ID: PAW-M08-U05 Estimated active time: 25-30 minutes
What you are producing
Two outputs for one task, both scored against the same rubric, with a decision recorded for each.
Generate two genuinely different versions
Use two different prompts for the same task — for instance your original and one improved with Module 6's example and rubric.
Or use the same prompt twice and accept whatever variation appears. Both are informative, for different reasons.
Score blind if you can
If someone else can label them A and B without telling you which prompt produced which, do that.
Knowing which one you expect to be better is a strong influence on scoring, and blind scoring is the only cheap way to remove it.
The comparison sheet
COMPARISON
Task:
Rubric criteria: 1. 2. 3. 4. 5.
Output A Output B
Accuracy
Completeness
Usefulness
Honesty
Tone
Decision keep/edit/... keep/edit/...
Time to review
What A had that B did not:
What B had that A did not:
Which prompt produced the better one, and what in it was responsible:
The last line is the exercise
Identifying *what in the prompt* was responsible is harder than picking a winner, and it is the only part that transfers to your next task.
If you cannot identify it, run a third version changing one thing at a time.
What to keep
Save the comparison sheet with both outputs.
Add whatever you identified as responsible to your improvement log, and to the saved prompt if it has now appeared more than once.
