Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

Bolt Performance Calibration — If Two Groups Start Different, the Final Difference May Not Be the Intervention

Wait, What? A Better Final Score Can Be Real and Still Overstate the Programme

A school pilots a new teaching approach in one set of classes and compares the results with another set. At the end of term, the intervention classes score 8 points higher.

That looks like an 8-point programme effect.

But suppose the intervention group already started 6 points ahead before the programme began. The final difference is still real. The causal story is now much less simple.

Before Bolt asks what improved, it asks whether the groups were comparable enough at the start for the later contrast to carry the interpretation the school wants to place on it.

Quick Answer

Owned Bolt job: calibrate intervention and school-improvement claims when the intervention and comparison groups begin at meaningfully different baseline levels or differ on important starting characteristics.

Baseline difference does not automatically prove a study is biased. In a properly randomised trial, some chance imbalance is expected and significance-testing of baseline differences is generally discouraged. In quasi-experimental comparisons, compromised randomisation, high-attrition analyses, or small cluster studies, however, baseline equivalence becomes much more important because starting differences may be entangled with the final outcome.

The practical rule is: do not let the final score erase the starting line.

Three Different Questions Are Hiding Inside “Which Group Did Better?”

1. Which group finished higher?

This is a descriptive question. If Group A finishes at 78 and Group B at 70, Group A finished higher.

2. Which group changed more?

This is a growth question. It requires trustworthy baseline and follow-up measures and a scale on which the change is interpretable.

3. Did the intervention cause the difference?

This is a causal question. It requires a credible counterfactual: what would have happened to the intervention students without the intervention?

A final mean can answer the first question while doing very little to answer the third.

Why Baseline Matters More in Some Designs Than Others

Proper random assignment

Randomisation is designed to make treatment assignment independent of pre-existing characteristics on average. With finite samples, the groups can still differ by chance. Modern trial guidance therefore recommends planning baseline adjustment in advance rather than deciding whether to adjust based on whether a baseline difference happens to be statistically significant.

Quasi-experimental comparison

If schools, teachers or students were not randomly assigned, starting differences are much more threatening. Intervention participants may already be more motivated, better resourced, higher attaining, more experienced, or selected precisely because they were struggling.

Randomised trial after substantial attrition

Randomisation protects the original groups. If many outcomes later go missing and the final analysis contains a different subset, the analysed groups can lose some of that original comparability. Baseline equivalence then becomes relevant again.

Small number of schools or classes

Randomising only a handful of schools or classrooms can leave large chance differences in prior attainment, leadership, cohort composition or teacher experience. Randomisation still matters, but a small number of assignment units creates less protection against inconvenient imbalance in any one realised trial.

The What Works Clearinghouse Makes Baseline Equivalence Explicit

The US Institute of Education Sciences’ What Works Clearinghouse standards briefs treat baseline equivalence as the degree to which intervention and comparison groups are similar at the start on characteristics expected to influence outcomes.

Under the current WWC framework, baseline equivalence is especially important for quasi-experimental studies and for randomised studies where attrition or randomisation integrity creates additional risk. The WWC handbook specifies formal thresholds and adjustment rules for studies seeking to meet its evidence standards. Those thresholds are not a universal school rule, but they demonstrate an important measurement principle: starting differences must be visible and, where necessary, accounted for rather than silently absorbed into the intervention effect.

WWC guidance also notes that even small low-attrition randomised trials can contain baseline differences by chance. The recommended response is not to declare the randomisation invalid simply because one baseline characteristic differs; it is to use design-consistent adjustment and report the uncertainty honestly.

School, Teacher and Student: Three Sources of Starting Difference

School

Schools can differ before an intervention starts: staffing stability, timetable allocation, prior attainment, leadership, resources, selection, attendance, subject uptake and other system conditions. If one school receives the programme and another acts as comparison, the school-level difference can masquerade as intervention effect.

Teacher or coach

Teachers who volunteer for coaching may differ from teachers who do not. They may be more motivated, more reflective, more experienced, or more concerned about current performance. A coaching outcome should therefore separate programme effect from self-selection where the design allows it.

Student

Students entering an intervention may differ in prior achievement, attendance, language background, exposure, support or motivation. These are not excuses and should not be turned into destiny. They are pre-intervention evidence that may influence later outcomes and therefore the interpretation of the comparison.

Competing Explanations for a Higher Intervention-Group Final Score

  • The intervention genuinely improved performance.
  • The intervention group already started stronger.
  • The comparison group contained students with greater initial need.
  • Teachers in the intervention group were selected because they were unusually motivated or experienced.
  • The programme interacted with baseline level, helping some starting profiles more than others.
  • The groups differed on another characteristic related to the outcome.
  • Attrition changed the analytical groups after assignment.
  • The final assessment itself favoured one group’s prior exposure or curriculum sequence.

Final-score difference alone cannot decide among these explanations.

The Bolt Baseline Calibration Protocol

  1. Freeze the comparison before looking at outcomes. Record assignment, eligibility and the intended analytical population.
  2. Name the primary outcome and relevant baseline measure. The strongest baseline covariate is usually closely related to the later outcome.
  3. Inspect starting distributions, not only means. Two groups can share an average while differing greatly in spread or lower-tail performance.
  4. Inspect assignment level. Student, class, teacher and school randomisation protect different levels of inference.
  5. Do not use baseline significance tests as a post-hoc switch in a proper RCT. Plan sensible covariate adjustment in advance.
  6. For non-randomised comparisons, establish equivalence before causal language becomes strong. Matching, weighting, regression adjustment or other defensible methods may help, but they cannot adjust perfectly for unmeasured differences.
  7. Check whether attrition changed the baseline profile of analysed participants.
  8. Report adjusted and unadjusted estimates where useful. A conclusion that survives both is more reassuring.
  9. Preserve uncertainty when groups remain meaningfully different. Narrow the causal claim instead of forcing certainty.
  10. Recalibrate the RFE decision. Decide whether evidence justifies scaling, continuing, redesigning, or collecting a stronger comparison.

Worked Example: The Stronger Class Gets the New Programme

A school gives a new science programme to Class A because its teacher volunteers. Class B continues with usual instruction. At baseline, Class A averages 68 and Class B 59. At follow-up, Class A averages 78 and Class B 67.

The final difference is 11 points. The raw gains are +10 and +8.

The result is consistent with a modest positive programme effect. It is also compatible with persistent baseline advantage and teacher self-selection. Without a stronger design or well-justified adjustment, “the programme caused an 11-point advantage” is too strong.

The school can still make an intelligent next decision: repeat the programme with more classes, use a stronger assignment or matching process, prespecify the analysis, and observe whether the advantage returns under a more credible comparison.

Why Statistical Adjustment Helps—but Does Not Time-Travel

Regression adjustment, ANCOVA, matching and weighting can improve comparability on observed baseline variables. They are valuable tools, and the WWC explicitly recognises several adjustment strategies in studies that need to establish baseline equivalence.

But adjustment cannot guarantee that unmeasured differences disappeared. If the intervention group had unusually strong leadership, parental support or teacher motivation that was never measured, no statistical model can recover that missing counterfactual perfectly.

This is why design quality comes before analytical cleverness. Better analysis can repair some imbalance. It cannot fully replace a credible comparison strategy.

What CONSORT 2025 Adds to the Calibration

The 2025 CONSORT explanation and elaboration makes an important distinction for randomised trials: baseline covariate adjustment should be specified because it improves precision or follows the design—not because investigators test baseline variables and decide after seeing which happened to be “significantly different.”

That principle transfers well to education. Baseline data should discipline interpretation, not become a post-hoc tool for manufacturing the preferred answer.

Common Misconceptions

  • “Any baseline difference invalidates a randomised trial.” False. Chance imbalance is expected in finite samples.
  • “If the baseline difference is not statistically significant, the groups are equivalent.” False. Non-significance does not prove meaningful equivalence.
  • “Adjustment removes all selection bias.” It addresses measured differences under assumptions; unmeasured confounding can remain.
  • “Only the final score matters.” Final status, growth and causal effect are different claims.
  • “A stronger baseline group should be penalised.” No. Baseline is used to interpret comparison, not morally discount achievement.
  • “Small trials cannot be useful.” They can be highly useful for learning and feasibility; they simply support narrower causal certainty.

How Do We Know?

The current What Works Clearinghouse standards resources define baseline equivalence and provide dedicated guidance on when intervention and comparison groups must demonstrate comparability. The WWC Procedures and Standards Handbook, Version 5.0 specifies acceptable baseline-adjustment approaches for quasi-experimental studies and randomised studies with relevant risk.

WWC’s Designing Quasi-Experiments guidance explains why baseline equivalence is central when treatment is not randomly assigned and why some starting differences require adjustment while larger differences can prevent a study from meeting WWC equivalence standards.

The 2025 CONSORT explanation and elaboration clarifies that in properly randomised trials, covariate adjustment should be prespecified and should not be triggered by significance tests of baseline imbalance.

Evidence boundary: formal thresholds and statistical procedures depend on the evaluation standard and design. A school should not mechanically import WWC cut-offs into every classroom comparison. The transferable principle is to make baseline comparability visible and match the causal claim to the design’s ability to separate starting difference from intervention effect.

For Parents: Ask Where the Groups Started

If a school says students in a new programme scored higher than other students, ask what the groups looked like before the programme. Were starting scores similar? Were students selected into the programme? Were the same kinds of learners represented?

This does not undermine success. It tells you whether the comparison is describing a promising programme or estimating its causal contribution.

Bolt RFE: What Should Change Next?

If baseline imbalance weakens the current interpretation, the next cycle should strengthen the comparison: improve assignment, prespecify adjustment, broaden the number of classes or schools, preserve baseline and follow-up data, and test whether the effect returns when starting conditions are better controlled.

A school should not scale because one group finished higher. It should scale when the evidence increasingly supports that the intervention—not merely the starting line—helped create the better finish.

Bolt Direction Graph

Intervention/comparison groups → baseline profile → assignment integrity → follow-up performance → adjusted/unadjusted comparison → inspect remaining imbalance → calibrated causal claim → strengthen next comparison → observe return → recalibrate.

Useful neighbours include If the Missing Students Are Different, the Intervention Effect Can Change, A Growth Score Can Be Noisier Than Either Test It Came From, and A Class Score Gain Is Not Automatically the Teacher’s Effect.