A fictional Python course has 120 students. Forty choose an optional practice session; thirty-two of them pass the final assessment. Of the eighty who do not choose it, forty pass.
The course team wants to advertise, “The practice session increases your chance of passing by thirty percentage points.” Write a claim this table could support, then name one difference between the groups that could explain the gap without a session effect. Explain why that difference would matter.
The observed pass rates are 80% and 50%, a thirty-percentage-point difference. That comparison describes this cohort. It does not, by itself, show that the session caused the difference. Students chose whether to attend, and the groups may have differed before the session.
The useful habit is to separate three questions: What happened in these records? Who else might this apply to? What caused it? Study design matters to the last two questions. OpenIntro's statistics text distinguishes sampling from a population and assigning a treatment in an experiment. Those choices support different kinds of conclusion. Study design
You will practise writing a conclusion that preserves the finding without claiming more than the data establish. No programming is needed.
Start with what one row represents
For this example, each row is one enrolled student, and every student's final result is recorded. The optional-session field records attendance; the result field records whether the student passed the same assessment.
| Optional practice session | Passed | Did not pass | Total | Pass rate |
|---|---|---|---|---|
| Attended | 32 | 8 | 40 | 80% |
| Did not attend | 40 | 40 | 80 | 50% |
| Entire cohort | 72 | 48 | 120 | 60% |
These totals let you check the denominator. The attended group has 32 ÷ 40 passes, not 32 ÷ 120. The cohort has 72 ÷ 120 passes, not the unweighted average of 80% and 50%. The groups have different sizes.
Now imagine the export has one row per page visit. A student who opens ten pages appears ten times. Counting rows would count visits, not students. Before computing a learner rate, establish which records belong to each learner and which outcome you intend to count.
Write the unit of observation beside your result: “120 students,” “1,200 page visits,” or “40 completed assessments” means different things. The label is part of the evidence.
State the comparison before explaining it
A supported description is:
In this cohort, 32 of 40 students who attended the optional session passed, compared with 40 of 80 who did not attend. The observed pass rates differed by thirty percentage points.
That statement preserves the useful pattern. It gives a teacher a reason to investigate the session and the students who chose it.
“Thirty percentage points” is an absolute difference: 80% − 50%. The relative difference is (80% − 50%) ÷ 50% = 60%. Saying “thirty percent higher” would blur those two calculations. Neither calculation supplies a causal explanation.
You do not have to remove every interesting finding because its interpretation has limits. You do have to keep the finding and its interpretation distinct.
Ask how the groups came to differ
Perhaps students already comfortable with Python were more likely to attend. Perhaps students with more available time attended and also studied longer at home. These are possible explanations to investigate, not facts established by the table.
An earlier knowledge check and a record of available study time could help examine those possibilities. They would not guarantee that every relevant difference had been measured. Observational comparisons can be affected by confounding: another factor can be related to both participation and the outcome. Random assignment in a suitably designed experiment helps address this problem; it differs from randomly selecting people to survey. OpenIntro on experiments and observational studies
Causal analysis can also use observational designs with explicit assumptions. This small table does not provide such a design. Avoid replacing its missing evidence with the word “proves.”
Check who the records can represent
The example includes all 120 enrolled students in one cohort, so it describes that cohort's recorded outcomes. It does not automatically represent every future learner, another course or students with different preparation.
To make a wider claim, ask how the students were recruited, what population you mean and whether the learning conditions match. A class of volunteers in an evening workshop may differ from students taking a required daytime course.
More rows from the same narrow source do not explain whether that source represents the wider group. Be specific about the population in your conclusion instead of treating “students” as an unlimited category.
Keep missing answers visible
Suppose the same course sends a recommendation survey to all 120 students. Thirty respond, and twenty-seven say they would recommend it.
“90% of respondents recommend the course” is supported: 27 ÷ 30 = 90%. “90% of enrolled students recommend the course” is not established. Ninety students have not answered.
The response rate is 30 ÷ 120 = 25%. That tells you how much response information you have; it does not measure the exact bias. Nonresponse bias depends on how respondents and nonrespondents differ on what you are measuring. The Federal Committee on Statistical Methodology cautions against using response rates alone to establish data quality. Nonresponse bias reporting
In this fictional example, the possible recommendation share among all enrolled students ranges from 27 ÷ 120 = 22.5%, if every nonrespondent would say no, to 117 ÷ 120 = 97.5%, if every nonrespondent would say yes. Those extreme bounds demonstrate what is unknown. They are not a confidence interval or an estimate of what the missing students believe.
Try a conclusion that stays within the evidence
The course team plans to use this statement to decide whether to expand the session:
Our optional practice session caused a thirty-point improvement, and 90% of our students recommend the course. The next class will get the same results.
Rewrite it in no more than three sentences, preserving the useful observed comparison and the survey counts. Then propose one way to gather stronger evidence for the session's effect and one way to learn about the missing survey opinions. State what each step would resolve and what uncertainty would remain.
A possible correction is:
In the fictional cohort, students who chose the optional session had an 80% pass rate, compared with 50% for those who did not. The comparison does not establish the session's causal effect or predict the next cohort's result. Twenty-seven of thirty survey respondents recommended the course; the opinions of ninety nonrespondents are unknown.
For the next evidence step, a suitably designed comparison with random assignment could help estimate the session's effect in that setting. Measuring earlier Python knowledge could investigate one alternative explanation, but it would not establish that all confounding is removed. Following up with nonrespondents could reduce the missing survey information; report how many still do not answer. Neither step guarantees that a future class will obtain the same results.
If your correction loses the observed difference, you have removed useful evidence. If it says the session “may have caused exactly thirty points,” you are still assigning a size to an effect the table has not identified. If your evidence plan only asks for “more data,” name which uncertainty those data would address.
Keep a short note with any analysis: the question, the unit represented by a row, the population covered, missing information, the calculation and the conclusion it supports. That note gives someone else a practical way to examine your reasoning.
Continue with Accuracy is not the whole evaluation to see how a correct percentage can still hide the mistakes that matter.