Module 9 Lesson: Comparing Groups and Effect Sizes
Classroom Explanation
Two questions get collapsed into one. Whether a difference is distinguishable from zero is a statistical question. Whether it is large enough to act on is a judgement about costs, and no calculation settles it.
In this module, the goal is to compare groups using magnitude, uncertainty, and practical meaning. The important part is not only the calculation. The important part is the thinking before and after the calculation.
Why This Matters
The common mistake is ranking groups by mean only and ignoring sample size, spread, and how people entered the group. This mistake can make a report sound more confident than the evidence deserves.
Key Ideas
1. A group comparison should name the groups, metric, and sample sizes.
A fair comparison names the groups, the metric, and how people or records entered each group. Without that, the comparison may be shallow.
2. Effect size asks how large the difference is.
Effect size asks how much difference there is. It moves the discussion from is there evidence? to how large is the difference?
3. A result can be statistically visible but practically small, or practically important but uncertain.
Practical meaning depends on the decision. A tiny change can be statistically visible and still not worth acting on.
Worked Example
The three-practice-quiz group has a higher mean score than the no-practice group. The optional review group is also high, but self-selection makes that comparison weaker.
Here is the dataset used in this module.
| group_name | sample_size | mean_score | standard_deviation | notes |
|---|---|---|---|---|
| No practice quiz | 48 | 68 | 12 | baseline group |
| One practice quiz | 52 | 72 | 11 | small improvement |
| Three practice quizzes | 50 | 79 | 10 | stronger improvement |
| Optional review session | 28 | 81 | 14 | self-selected group |
| Required review session | 55 | 76 | 9 | assigned group |
Decide the smallest difference worth acting on before analysing anything. With a large enough sample almost any non-zero difference becomes statistically detectable, so without a threshold agreed in advance the answer will always turn out to be just above whatever you observed.
How To Think Through It
- State the difference in the units of the measure.
- Report its interval.
- Compare the whole interval with your pre-set threshold.
- Say which of the four possible verdicts applies.
- Recommend, including the option of doing nothing.
Common Mistake
The common mistake is ranking groups by mean only and ignoring sample size, spread, and how people entered the group.
How To Write The Result
The three-practice-quiz group shows a larger average score than the no-practice group, but the comparison should still mention spread and assignment conditions.
Practice Prompt
Write a cautious group-comparison paragraph using mean, sample size, spread, and one limitation.
Takeaway
Statistical significance and practical significance are different findings; report both.
