Wait, What? A Programme Can “Work” on One Measure Even When the Overall Evidence Is Weak
A school evaluates a new programme using reading score, writing score, attendance, confidence, homework completion, classroom behaviour, motivation, quiz performance, examination performance, teacher rating and student survey.
Ten outcomes are tested. Nine show little or no clear difference. One shows a positive result.
The school announces: “The programme significantly improved student motivation.”
That result may be real. But when many outcomes are examined, the chance of finding at least one apparently impressive result rises even if the intervention has no broad effect.
Quick Answer
Owned Bolt job: calibrate intervention-effect claims when a school measures many outcomes, time points, subgroups or analyses and highlights only the most favourable result.
Multiple outcomes are often necessary because education is multidimensional. The problem is not measurement breadth. The problem is selective interpretation after the results are known. Strong evaluation distinguishes primary outcomes from secondary or exploratory outcomes, prespecifies important analyses, reports the full pattern, and accounts for multiplicity where appropriate.
The RFE question is simple: Did the programme improve the educational target we cared about before seeing the data, or did we find a favourable result only after searching across many possibilities?
Why More Outcomes Create More Opportunities for False Confidence
Every statistical test has some chance of producing an apparently positive result even when the underlying effect is absent. When a study runs many tests, those opportunities accumulate.
This can happen through:
- many outcome measures;
- many subscales inside one test;
- several follow-up time points;
- many student subgroups;
- multiple treatment groups;
- several alternative statistical models;
- many ways of defining “success.”
If only the most favourable result is reported, readers see a cleaner story than the evidence actually produced.
Primary, Secondary and Exploratory Outcomes Should Not Be Treated as Interchangeable
Primary outcome
The main outcome the evaluation is designed to answer. It should usually be chosen before results are known and should align with the intervention’s central educational job.
Secondary outcomes
Important additional outcomes that help explain breadth, mechanism or unintended effects.
Exploratory outcomes
Analyses used to generate hypotheses or investigate interesting patterns that were not the main confirmatory question.
All three are useful. They simply carry different evidential weight.
If a tutoring programme was designed primarily to improve mathematics achievement but shows no clear mathematics effect and one positive confidence result among many secondary outcomes, the school should not quietly redefine the programme as a proven confidence intervention without replication.
The What Works Clearinghouse Explicitly Corrects for Multiple Comparisons
The US Institute of Education Sciences’ What Works Clearinghouse has long treated multiple comparisons as a source of overstated statistical significance. Its standards note that when a study examines many outcomes or groups, the likelihood of finding an apparently significant result rises. The WWC therefore applies multiple-comparison adjustments within outcome domains when needed.
This is not an obscure statistical preference. It protects a basic educational question: are we seeing a robust intervention signal, or one favourable result selected from a large field of tests?
CONSORT 2025 Makes Selective Outcome Reporting Visible
The 2025 CONSORT reporting standard states that primary and secondary outcomes should be prespecified and completely defined. It also requires important changes to outcomes or analyses after a trial begins to be reported with reasons.
The reason is straightforward: research has repeatedly found discrepancies between outcomes specified in protocols or registries and outcomes highlighted in final reports, often in favour of statistically significant results.
Education can learn directly from this discipline. A school pilot may not be a journal trial, but it should still know which outcome mattered most before the dashboard lights up.
School, Teacher and Student: Three Different Ways Selective Success Appears
School
Schools can unintentionally promote the metric that looks best after implementation: attendance if grades are flat, confidence if attainment is flat, one subject if the others are flat, or one subgroup if the whole cohort is flat. This can make a mixed result look uniformly successful.
Teacher or coach
A coach may try several instructional changes and remember the one lesson where student response looked unusually strong. Without a defined target and repeated return evidence, coaching can become story-selection rather than calibration.
Student
A student can also cherry-pick performance evidence: “I got 90% on one quiz” while ignoring several weaker assessments. Bolt does not erase the good result. It places it inside the full evidence pattern.
Competing Explanations for One Positive Result Among Many
- The intervention genuinely affects that one outcome and not the others.
- The result is a chance false positive.
- The primary outcome measure was insensitive while the secondary measure captured a real mechanism.
- The positive result reflects a subgroup effect rather than a whole-cohort effect.
- The analysis was selected because it looked favourable after the data were inspected.
- The intervention needs longer for the primary outcome but affects an earlier intermediate outcome first.
- The measure is noisy and the apparent result will not replicate.
These explanations imply very different next decisions. One positive p-value cannot choose among them.
The Bolt Multiple-Outcomes Calibration Protocol
- Name the primary outcome before results are known. Align it to the intervention’s main educational job.
- Define secondary outcomes in advance. They can broaden interpretation without replacing the primary question post hoc.
- Record analysis timing. Distinguish prespecified from exploratory analyses.
- Report the full pattern. Positive, null and negative findings all belong to the evidence.
- Account for multiplicity when many related tests are used. The exact method depends on design and purpose.
- Prefer effect size and precision to significance alone. A tiny uncertain effect should not become a large narrative because it crosses a threshold.
- Treat subgroup findings cautiously unless prespecified and well powered.
- Replicate surprising exploratory findings. A second independent receipt is especially important when the result emerged from many searches.
- Check mechanism coherence. Does the positive outcome fit a plausible chain from teaching change to student performance?
- Recalibrate the RFE decision. Decide whether to scale, continue testing, narrow the claim, or redesign measurement.
Worked Example: Twelve Outcomes, One Improvement
A school introduces a study-support programme. It prespecifies mathematics attainment as the primary outcome and tracks eleven secondary outcomes including attendance, homework completion, confidence, stress, participation and several survey subscales.
Mathematics attainment shows little difference. Ten secondary outcomes also show little difference. One confidence subscale improves.
The school has three honest options:
- conclude that the programme did not yet demonstrate its intended mathematics effect;
- report the confidence result as exploratory evidence worth investigating;
- run a follow-up study designed specifically to test whether the confidence effect replicates and matters educationally.
The dishonest option is to advertise “proven confidence improvement” while making the original primary question disappear.
Prespecification Does Not Mean Researchers Must Stop Thinking
Exploration is essential. Unexpected results create new knowledge. The problem is not changing one’s mind; it is presenting an after-the-fact discovery as though it had been the planned confirmatory test all along.
EEF’s current evaluation infrastructure reflects this distinction. Its evaluator resources include protocol templates and statistical analysis plan templates, and current trials publish protocols and analysis plans before final results. That makes later deviations inspectable rather than invisible.
Schools can adopt the same principle without bureaucracy: write down the main outcome, main comparison and main decision rule before opening the final spreadsheet.
Why “Statistically Significant” Is Not the Same as “Educationally Important”
Multiple-comparison problems are often discussed through statistical significance, but Bolt needs a wider calibration.
A result can be statistically detectable and educationally trivial. Another can be educationally meaningful but estimated too imprecisely to cross a conventional threshold in a small pilot. Schools therefore need the magnitude of change, confidence around the estimate, implementation evidence, and consequences for actual student performance—not only a binary significant/not-significant label.
Common Misconceptions
- “Measuring many outcomes is bad.” False. Breadth can reveal mechanisms and unintended effects.
- “Only the primary outcome matters.” Secondary and exploratory outcomes are valuable; they simply support different strengths of claim.
- “A multiple-comparison correction tells us whether the intervention is educationally useful.” It addresses one statistical problem, not implementation, validity or practical importance.
- “If one result is significant, the programme worked.” The result must be interpreted inside the full outcome family and design.
- “If the primary outcome is null, every secondary finding is meaningless.” No. Secondary findings can generate important hypotheses and sometimes reflect real differentiated effects.
- “Prespecification prevents adaptation.” It prevents hidden rewriting of the original question; transparent changes remain possible.
How Do We Know?
The What Works Clearinghouse standards handbook explains that studies examining many outcomes or groups can overstate statistical significance if multiplicity is ignored and describes the WWC’s use of multiple-comparison adjustments.
WWC’s current developer and researcher guidance likewise explains that WWC reviews can alter reported significance after adjusting for multiple comparisons within an outcome domain.
The 2025 CONSORT explanation and elaboration states that trials should prespecify primary and secondary outcomes, disclose outcome changes, report planned outcomes rather than only interesting significant results, and address multiplicity when many analyses are conducted.
The Education Endowment Foundation’s current protocol, study-plan and statistical-analysis-plan templates provide a practical education-sector model for making the main evaluation questions and analyses visible before final outcome interpretation.
Evidence boundary: there is no single universal correction appropriate to every collection of outcomes. Multiplicity strategy depends on whether outcomes are confirmatory, secondary or exploratory, their correlation, the design and the decision being made.
For Parents: Ask What the Programme Was Supposed to Improve Before the Results Arrived
If a programme advertises one strong result, ask what its primary goal was and what happened to the other outcomes measured. A trustworthy evaluation should be able to explain the whole pattern rather than only its most photogenic number.
Bolt RFE: What Should Change Next?
If one favourable outcome appears inside an otherwise weak evidence pattern, the next cycle should not immediately scale the programme. Prespecify the new hypothesis, measure it directly, repeat the intervention, and see whether the result returns under an evaluation designed to test that effect rather than discover it accidentally.
The strongest school is not the one that can find a success metric. It is the one whose success claim survives the outcomes it hoped would work, the outcomes that did not, and the next independent test.
Bolt Direction Graph
Intervention job → prespecified primary/secondary outcomes → many observed results → multiplicity + full-pattern check → classify confirmatory versus exploratory evidence → replicate surprising findings → recalibrate programme claim → next performance cycle.
Useful neighbours include If Two Groups Start Different, the Final Difference May Not Be the Intervention, You Cannot Judge an Intervention Before You Know What Was Actually Implemented, and One Result Should Not Rewrite the Whole Model.
