Wait, What? “After” Is Not the Same as “Because Of”
A school introduces a new programme in January. Students score 60 on a baseline assessment and 72 in June.
The gain is real: twelve points.
But students were also five months older. They received ordinary teaching, homework, revision, tutoring, peer support, school activities and examination practice. The curriculum moved forward. The baseline test itself may have taught them something about what to expect.
Without a credible comparison, the school knows that performance changed after the programme. It does not yet know how much of that change the programme caused.
Quick Answer
Owned Bolt job: calibrate school, teacher and intervention claims when performance is measured before and after a programme but there is no credible comparison group or counterfactual.
Pre–post evidence is useful for detecting change. By itself, it is weak for determining causation because many things besides the intervention can change over the same period. The correct RFE response is not to discard the gain. It is to separate observed improvement from programme-attributable improvement and design the next cycle so those quantities become easier to distinguish.
A Pre–Post Study Answers One Question Very Well
If the same students are measured before and after, a pre–post design can tell us:
Did measured performance change over this period?
That can be extremely valuable for school improvement. It tells teachers whether the starting state and later state differ.
What it does not identify cleanly is:
How much of that change would not have happened without the programme?
That is the counterfactual problem.
Five Things Can Change Between Baseline and Follow-Up Even If the Programme Does Nothing
1. Maturation
Students develop with age, experience and ordinary schooling. Younger learners in particular may improve on many tasks simply because time and normal education passed.
2. History
Other events occur during the intervention window: a new teacher, timetable change, examination revision, school policy, tutoring, holiday, disruption or subject-specific campaign.
3. Testing or familiarity effects
The first assessment can change later performance by familiarising students with content, format or expectations.
4. Regression toward typical performance
If students were selected because they had unusually low results, some improvement may occur simply because an extreme first result is followed by a more typical one.
5. Measurement change
The later test may differ in difficulty, content coverage, scoring, rater severity or administration conditions.
The programme can still be effective. These are competing explanations for the observed gain.
The Institute of Education Sciences Gives a Classic Warning
An IES evidence guide uses an unusually clear example: if researchers had evaluated an early-childhood programme only with a pre–post design, they could have concluded that the programme improved school readiness because participants improved over time. The randomised comparison showed that control-group children improved similarly, indicating that much of the apparent gain would have occurred anyway.
The same guide gives the reverse example. In a summer programme, participant reading performance declined. A pre–post design could have labelled the programme harmful. The control group declined even more, meaning the programme was actually protective relative to what would otherwise have happened.
This is the central Bolt lesson: without a counterfactual, even the direction of the programme effect can sometimes be misread.
A Modern EEF Pilot Makes the Same Point
In its evaluation of the RETAIN professional-development pilot, the Education Endowment Foundation reported increases in early-career teachers’ knowledge, confidence and classroom practice. But the evaluation explicitly stated that the absence of a comparison group meant it was not possible to estimate how much improvement might have occurred anyway through maturation and ordinary school support.
That is high-integrity reporting. The improvement is not denied. The causal claim is simply kept inside the evidence boundary.
School, Teacher and Student: Three Versions of “Would It Have Happened Anyway?”
School
A school may introduce a new curriculum, platform, coaching model or attendance intervention while many other system changes occur. If results improve, leaders should ask which changes were unique to the intervention group and which affected everyone.
Teacher or coach
A teacher changes questioning practice and later sees stronger student responses. That is promising. But the class may also have become more familiar with the topic, stronger in prerequisites, or more comfortable with discussion. Repeating the practice in another class or using staggered introduction can strengthen attribution.
Student
A student begins tutoring and improves over a term. The tutoring may have helped greatly. The student also received school teaching, revision and more practice. The honest learner model can say “performance improved during tutoring” before saying exactly how much tutoring caused the gain.
Competing Explanations for a Strong Pre–Post Gain
- The programme genuinely caused substantial improvement.
- Students would have improved through ordinary teaching anyway.
- The first test was unusually low.
- The later test was easier or more familiar.
- Another school-wide change improved performance.
- The programme accelerated improvement that would otherwise have happened more slowly.
- The programme prevented a decline that the pre–post scores do not reveal.
- Several causes contributed simultaneously.
A no-comparison design cannot distinguish these confidently from the two scores alone.
The Bolt Counterfactual Calibration Protocol
- Name the observed change. Keep the descriptive gain separate from the causal claim.
- Map everything else that changed. Teaching, timetable, staffing, assessment, revision, attendance and external support can matter.
- Ask what would normally happen over the same period. Historical cohorts, benchmarking or external data can provide context, though weaker than a strong concurrent comparison.
- Add a comparison group where feasible. Random assignment is strongest when appropriate; matched or otherwise credible comparisons can still improve the counterfactual.
- Use staggered implementation when ethical and practical. Different start times can create additional comparison opportunities.
- Use repeated pre-intervention observations where possible. A prior trajectory helps distinguish programme change from an existing trend.
- Keep assessments comparable. Do not confuse intervention effect with test change.
- Check whether the result survives fresh, delayed performance. This strengthens the learning claim even when causal attribution remains uncertain.
- Do not promote mechanism evidence into causal proof. Better implementation and plausible student response help the story but do not replace a counterfactual.
- Recalibrate the RFE decision. Decide whether evidence supports scaling, a stronger trial, continued local use, or only a promising hypothesis.
Worked Example: Everyone Improves After the New Revision Programme
A school introduces a revision programme to Secondary students eight weeks before examinations. Mean scores rise from 64 to 74.
The programme looks powerful. But revision intensity also increased across every subject, students completed several practice papers, and the follow-up test occurred closer to the examination after substantial ordinary teaching.
The next year, the school introduces the programme to half the classes first and delays it for the others by four weeks. During the initial comparison window, early-start classes improve faster on fresh comparable assessments. After the delayed classes receive the programme, their performance also rises.
The second design does not make every confound disappear, but it produces a much stronger causal story than the original pre–post gain alone.
Sometimes a Comparison Shows That a Programme Helped Even When Scores Fell
This is one of the most counterintuitive reasons counterfactuals matter.
Suppose examination performance normally drops after a long holiday. Students in a summer programme fall by 3 points, so a simple pre–post report says the programme failed. Comparable students without the programme fall by 9 points.
The intervention group’s score still declined. Relative to the credible counterfactual, the programme may have prevented a larger loss.
Performance calibration therefore cannot be reduced to “up is good, down is bad.” The missing question is always: up or down compared with what would probably have happened otherwise?
What About Interrupted Time Series?
If randomised comparison is impossible, repeated measurements across a long pre-intervention and post-intervention period can improve inference. Interrupted time-series approaches ask whether the level or trend changes when the intervention begins.
Comparative interrupted time-series designs go further by adding a comparison series. IES has funded methodological work testing when these designs can approximate credible programme effects in education.
These designs are more informative than a single before/after pair, but they still rely on assumptions about trends, concurrent events and comparison quality. They are tools for improving the counterfactual, not machines that manufacture certainty.
Common Misconceptions
- “Pre–post studies are useless.” False. They are excellent for describing change, feasibility and local signals.
- “If the gain is huge, a control group is unnecessary.” Large gains can still reflect maturation, regression, testing or concurrent change.
- “A comparison group must receive nothing.” It can receive usual practice or another active condition, depending on the question.
- “Randomised trials are always possible.” Practical and ethical constraints matter; stronger quasi-experimental designs can still improve inference.
- “A control group proves causation automatically.” Attrition, contamination, baseline problems and implementation can still weaken the comparison.
- “If scores fall, the intervention harmed students.” Not without knowing the expected trajectory without intervention.
How Do We Know?
The IES guide Identifying and Implementing Educational Practices Supported by Rigorous Evidence explains directly why pre–post studies cannot determine whether observed improvement or decline would have happened anyway without the intervention, and provides education examples where a pre–post interpretation would have reached the wrong causal conclusion.
The Education Endowment Foundation’s RETAIN pilot evaluation reports improvements in teacher knowledge and practice but explicitly states that, without a comparison group, it was not possible to estimate how much improvement might have occurred through maturation and ordinary school support.
IES’s project on Robustness of Comparative Interrupted Time Series Designs in Practice examines when stronger longitudinal comparison designs can produce trustworthy education programme-effect estimates when conventional randomised designs are unavailable.
The current What Works Clearinghouse evidence system continues to privilege designs that make the comparison and counterfactual explicit when determining whether an intervention has strong, moderate or promising evidence of effectiveness.
Evidence boundary: a comparison group is only as useful as its credibility. Poorly matched controls can mislead, and randomisation can be weakened by attrition, contamination or implementation problems. The general principle is not “always demand an RCT”; it is “match causal confidence to how well the design answers what would have happened otherwise.”
For Parents: Improvement After Tuition Is Good News—Then Ask What the Evidence Can Actually Attribute
If a child improves after tuition, coaching or a school programme, the improvement is worth celebrating. The causal question is separate. Was the child also receiving more revision? Did the examination season approach? Did school teaching change? Does the improvement persist on fresh tasks and after support is reduced?
You do not need to solve the counterfactual perfectly to make a sensible decision. You do need to avoid pretending that timing alone proves cause.
Bolt RFE: What Should Change Next?
If the current evidence is only pre–post, use it as a strong signal for the next evaluation rather than the final verdict. Add a comparison, stagger implementation, preserve repeated baseline observations, specify the primary outcome, and test whether the improvement still appears relative to what comparable learners would otherwise have done.
The purpose of a counterfactual is not to make improvement harder to believe. It is to discover which change deserves to be repeated because it actually helped create the better performance.
Bolt Direction Graph
Baseline performance → programme begins → follow-up performance → observed pre–post change → map maturation/history/testing/regression → add credible counterfactual → estimate attributable change → choose next intervention decision → observe return → recalibrate.
Useful neighbours include A Growth Score Can Be Noisier Than Either Test It Came From, If Two Groups Start Different, the Final Difference May Not Be the Intervention, and You Cannot Judge an Intervention Before You Know What Was Actually Implemented.
