Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

Bolt Performance Calibration — A Statistically Significant Result Can Still Be Too Small to Matter Educationally

Wait, What? A Result Can Be “Statistically Significant” and Still Be Too Small to Change What a School Should Do

A school trials a new programme with hundreds of students. The analysis reports a statistically significant improvement. The headline sounds decisive: the intervention worked.

But suppose the average gain is tiny, expensive to produce, disappears six weeks later, does not transfer to ordinary classroom performance and is smaller than normal score uncertainty for many individual students. The statistical result may still be correct. The educational interpretation may still be weak.

Bolt separates two questions that are often collapsed: Is there evidence that the observed difference is unlikely to be explained by sampling variation alone? And is the estimated difference large, durable, transferable and useful enough to matter educationally?

Quick Answer

Owned Bolt calibration job: distinguish statistical detectability from educational importance when interpreting intervention, programme or teaching effects, so a small p-value is not automatically converted into a large claim about learning.

Statistical significance answers a limited question about the compatibility of observed data with a statistical model and null hypothesis. Educational importance requires additional evidence: magnitude, uncertainty, baseline performance, outcome quality, duration, transfer, implementation burden, cost, equity and the decision being made.

A school can therefore have a statistically significant result that is educationally trivial, and an educationally promising effect that fails to reach statistical significance in a small or noisy study. Neither possibility licenses wishful thinking. Both require calibrated interpretation.

Statistical Significance and Effect Size Do Different Jobs

A p-value is influenced by the size of the observed effect, the amount of variability and the amount of information in the study. With a very large sample, a very small average difference can be detected precisely enough to cross a conventional significance threshold. With a small sample, even a practically important effect can remain statistically uncertain.

An effect size describes the magnitude of a difference or relationship on a defined scale. It does not tell us automatically whether that magnitude matters. Educational importance depends on the context in which the effect occurs.

Matthew Kraft’s education-specific work on interpreting intervention effects makes this point directly: generic labels such as “small,” “medium” and “large” can mislead because the size of typical field-based education effects, the cost of producing them and the difficulty of scaling them differ from one intervention and setting to another.

Observable Evidence Pattern

Study resultPossible statistical readingBolt calibration question
Very small effect, huge sample, p < .001Difference is estimated preciselyIs the magnitude educationally meaningful?
Moderate effect, small sample, p = .08Evidence remains uncertainIs more evidence needed before rejecting a potentially useful intervention?
Significant immediate gain, no delayed gainShort-term effect detectedDid learning persist?
Significant researcher-made test gain, no independent-test gainEffect depends on outcome measureDid the intervention improve broader performance or mainly the aligned measure?

The table shows why “significant/not significant” is too low-resolution for educational decision making.

What Statistical Significance Can Support

When the study design and analysis are appropriate, statistical significance can support the claim that the observed result is difficult to reconcile with a specified null model at the chosen threshold. It can be useful evidence that a difference deserves attention.

It is especially useful when combined with an estimated effect, confidence interval, transparent outcome definition and a credible design. The signal is stronger when the effect replicates across studies, settings or independent outcomes.

What Statistical Significance Cannot Support by Itself

  • It cannot tell us that the effect is large.
  • It cannot tell us that the effect is important to students, teachers or schools.
  • It cannot establish that the intervention caused the effect if the design does not support causal inference.
  • It cannot show whether the gain lasts.
  • It cannot show whether the gain transfers beyond the measured outcome.
  • It cannot tell us whether the programme is worth its cost, time or opportunity cost.
  • It cannot tell us whether benefits and burdens are distributed fairly.
  • It cannot turn a group-average effect into a guaranteed individual benefit.

The Outcome Measure Can Change the Apparent Effect

An intervention may produce a larger effect on a researcher-designed measure that closely matches the taught material than on an independent assessment. This does not automatically make the researcher measure invalid. It does mean that the magnitude of the effect depends partly on what was measured.

A What Works Clearinghouse working paper using WWC study data found that effect sizes on researcher- or developer-created outcomes were substantially larger on average than effects on independent measures. That is a direct warning against treating an effect size as a property of the intervention alone.

Bolt therefore asks: effect on what? A statistically significant improvement on a narrow near-transfer measure should not silently become “students became broadly better learners.”

Educational Importance Needs a Decision Frame

The same effect can matter differently in different settings. A very small average improvement may be valuable if the intervention is almost free, reaches millions of learners, reduces an important inequality and has no meaningful downside. A larger effect may be unattractive if it requires unsustainable staffing, displaces more effective teaching time or disappears as soon as external support is removed.

That is why education-specific effect-size interpretation should consider study features, costs and scalability rather than importing universal thresholds mechanically.

School–Teacher–Student Triad

School

The school should ask whether the detected effect is large enough, reliable enough and relevant enough to justify changing policy, curriculum, staffing or spending. It should inspect the confidence interval, outcome measure, implementation fidelity, subgroup pattern, cost and persistence—not only whether p crossed .05.

Teacher or Coach

The teacher should translate group evidence cautiously. If a programme has a modest average effect, that does not predict that every student will benefit. The teacher still needs local performance receipts: what changed for this learner, under what conditions, and did the change survive later independence?

Student

The student should not hear “research says it works” as a command to ignore their own performance evidence. Research can justify trying an intervention. The learner’s later performance helps decide whether it is working for this learner in this context.

The Bolt Significance-to-Importance Calibration Protocol

  1. Name the causal question. What intervention or condition is being compared with what alternative?
  2. Check the design. Does the study support causal, correlational or merely descriptive claims?
  3. Read the effect estimate. Do not stop at the p-value.
  4. Read the uncertainty interval. What range of effects remains compatible with the evidence?
  5. Inspect the outcome. Is it independent, curriculum-aligned, near-transfer, delayed or long-term?
  6. Check durability. Does the gain survive after the immediate intervention period?
  7. Check transfer. Does it appear on a different task or only the trained measure?
  8. Check burden and cost. What did the effect require from teachers, students and the school?
  9. Check distribution. Did some learners benefit while others did not?
  10. Make the smallest justified decision. Pilot, continue, stop, adapt or collect more evidence.

Worked Example: The Programme That “Worked”

A school introduces a digital mathematics programme for 1,500 students. The average end-of-term standardised score is 0.06 standard deviations higher than the comparison group, and the difference is statistically significant.

A headline could say, “Programme produces significant mathematics gains.” Bolt asks more. The programme costs substantial lesson time and licensing fees. The effect is concentrated on a platform-aligned test. On a common school examination, the difference is 0.01 standard deviations with a wide interval. Six weeks later, the groups are indistinguishable.

The calibrated conclusion becomes: the programme produced a small detectable near-term effect on the aligned outcome, but current evidence does not justify a strong claim of durable, transferable mathematics improvement.

That conclusion does not deny the statistical result. It gives the statistical result the correct educational size.

Prediction Before the Next Cycle

If the intervention genuinely improves broad capability, school and teacher should be able to predict where the gain will reappear: perhaps on an independent delayed assessment, a transfer problem or a reduction in the lower tail of performance.

The next performance becomes a receipt. If the predicted transfer fails repeatedly, the school should update its interpretation even if the original p-value remains unchanged forever.

Teacher–Student Dialogue

Student: “The programme is proven to work, right?”

Teacher: “The study detected an average difference. That is useful evidence. We still need to know how large it is, what it improved and whether that improvement lasts.”

Student: “So research can be right without the programme being worth it?”

Teacher: “Exactly. Statistical evidence and educational value are related, but they are not the same judgement.”

For Parents

When a school says an intervention is “statistically significant” or “evidence-based,” ask what changed and by how much. Ask whether the outcome was an independent assessment, whether gains persisted, what the programme costs in time and money, and whether children similar to yours benefited.

A statistically significant result is not fake or meaningless. It is simply not the final educational decision.

How Do We Know?

Matthew Kraft’s 2020 article Interpreting Effect Sizes of Education Interventions argues that generic effect-size benchmarks are often poor guides for education and proposes interpretation that considers study features, costs and scalability.

The What Works Clearinghouse Procedures and Standards Handbook separates study quality, effect estimates and statistical inference rather than treating one significance test as the entire evidence judgment. WWC resources also report effect magnitude alongside statistical significance for intervention findings.

The WWC working paper Average Differences in Effect Sizes by Outcome Measure Type found that effect sizes based on researcher- and developer-created measures were substantially larger on average than effects based on independent measures. This directly supports checking what outcome produced the effect before interpreting its educational reach.

Earlier WWC standards explicitly distinguished statistical significance from a separate concept of substantive importance, illustrating why detectability and magnitude are different dimensions even though current WWC characterisation rules have evolved.

Evidence Boundary

There is no universal effect-size threshold below which an educational intervention never matters or above which it always does. Outcome scale, age, subject, baseline risk, cost, duration, implementation and distribution all change interpretation. Bolt therefore does not invent a proprietary “meaningful effect” cut-off.

Statistical significance also does not diagnose why a particular student improved or declined. Individual educational observations remain non-clinical and should not be used to diagnose health or developmental conditions.

Common Misconceptions

  • “p < .05 means the effect is important.” It does not measure educational importance.
  • “A non-significant result means the intervention does nothing.” The study may be too small or noisy to estimate the effect precisely.
  • “A larger effect size is always better.” Cost, outcome quality, scalability and durability matter.
  • “If one measure improved, learning improved broadly.” Transfer must be demonstrated, not assumed.
  • “Research evidence replaces local evidence.” Research informs the trial; later learner performance helps calibrate the local decision.

What Should Change Next?

Suppose Bolt concludes: “The intervention has a small but credible average effect, and the next uncertainty is whether the claimed improvement appears in independent or transfer performance.” The next Bolt move is to collect that decision-matched receipt under declared conditions before enlarging the claim.

If that receipt exposes a genuine learning need, Bolt records the narrower learner-performance finding without prescribing the learner operation. The intervention claim and the learner-level claim remain separate evidence objects.

RFE: Did the later independent or transfer performance justify treating the detected average effect as educationally important for this decision, or should the school narrow, adapt or withdraw the original claim?

Bolt Direction Graph

Study design → effect estimate + uncertainty → statistical significance → magnitude/outcome/durability/cost check → calibrated educational-importance claim → later independent/transfer performance where needed → school/teacher/student recalibration.

Useful neighbours: Bolt — If the Missing Students Are Different, the Intervention Effect Can Change, Bolt — If the Control Group Starts Copying the Intervention, the Comparison Stops Being Clean, and Bolt — If You Measure Enough Outcomes, One Success Can Appear by Chance.