Module 10 Worksheet: Non-Parametric Thinking and Assumption Caution
Scenario
The values are strongly skewed and the summary you were given uses the mean.
Data
Dataset: ordinal-satisfaction.csv
| response_id | group_name | satisfaction_rank | completion_time_minutes | unusual_flag |
|---|---|---|---|---|
| O01 | Original | 3 | 18 | no |
| O02 | Original | 2 | 22 | no |
| O03 | Original | 4 | 19 | no |
| O04 | Original | 3 | 21 | no |
| O05 | Original | 1 | 35 | very slow completion |
| O06 | Original | 3 | 20 | no |
Your Task
Identify why a standard mean comparison may be weak for satisfaction ranks.
Questions
- Why is the mean a poor summary of this set?
- What do the median and the quartiles say instead?
- Which values are driving the mean?
- Should they be excluded, investigated, or kept?
- When would ranking be more honest than averaging?
- What summary would you publish?
- What does a rank-based summary not tell you?
Write Your Conclusion
Report a rank-based summary and say explicitly what it does and does not support.
Self-Check
- Did I identify the values driving the mean?
- Did I use ranks rather than magnitudes where appropriate?
- Did I investigate rather than delete?
- Did I say what the rank summary cannot tell me?
