Skip to course content
Free prompt workflow course

Prompting and AI Workflow Design

Practice: Evaluate Two Outputs for the Same Task

Unit ID: PAW-M08-U05 Estimated active time: 25-30 minutes

What you are producing

Two outputs for one task, both scored against the same rubric, with a decision recorded for each.

Generate two genuinely different versions

Use two different prompts for the same task — for instance your original and one improved with Module 6's example and rubric.

Or use the same prompt twice and accept whatever variation appears. Both are informative, for different reasons.

Score blind if you can

If someone else can label them A and B without telling you which prompt produced which, do that.

Knowing which one you expect to be better is a strong influence on scoring, and blind scoring is the only cheap way to remove it.

The comparison sheet

COMPARISON

Task:
Rubric criteria:  1.  2.  3.  4.  5.

              Output A          Output B
Accuracy
Completeness
Usefulness
Honesty
Tone

Decision      keep/edit/...     keep/edit/...
Time to review
What A had that B did not:
What B had that A did not:
Which prompt produced the better one, and what in it was responsible:

The last line is the exercise

Identifying *what in the prompt* was responsible is harder than picking a winner, and it is the only part that transfers to your next task.

If you cannot identify it, run a third version changing one thing at a time.

What to keep

Save the comparison sheet with both outputs.

Add whatever you identified as responsible to your improvement log, and to the saved prompt if it has now appeared more than once.