Wait, What? “It Didn’t Work” May Mean “We Never Really Ran It”
A school introduces a new teaching approach. Six months later, student results have barely moved. The programme is declared ineffective.
But some teachers used the approach weekly, some occasionally, some changed major components, and some never had the materials or time assumed by the design. In that situation, the school has not yet cleanly measured the intervention’s effect. It has measured the outcome of a mixed implementation.
Quick Answer
Owned Bolt job: calibrate school-performance claims by distinguishing intervention failure from implementation failure.
Before concluding that a teaching programme, coaching model, curriculum routine or school intervention “worked” or “failed,” schools should know what was actually delivered, to whom, with what dosage, quality, adaptation and support. Implementation fidelity is not administrative trivia. It is part of the evidence needed to interpret outcomes.
The Intervention on Paper and the Intervention in Classrooms Are Not the Same Object
Every school initiative has at least two versions:
- Intended intervention: what the design says should happen.
- Enacted intervention: what teachers and students actually experienced.
If those versions differ substantially, outcome interpretation becomes harder. A weak result could mean the idea itself was ineffective. It could also mean the intervention arrived incompletely, inconsistently, too late, without adequate training, or in a form that no longer preserved its active ingredients.
Implementation Fidelity Is More Than “Did Teachers Use It?”
A useful implementation record may include several dimensions:
- Adherence: were the intended components delivered?
- Dosage: how often and for how long?
- Quality: was the practice enacted competently?
- Reach: which students or classes actually received it?
- Responsiveness: did teachers and students engage with it?
- Adaptation: what was changed locally, and did the change preserve the function?
- Support conditions: were training, materials, leadership time and follow-up available?
Two schools can both report “we implemented the programme” while producing very different educational realities.
Fidelity Does Not Mean Blind Obedience
Implementation fidelity is often misunderstood as rigidly copying every detail. Good implementation research is more sophisticated. Some components may be essential; others may be adaptable. Context matters. A school may need to change scheduling, examples, language, materials or sequencing while preserving the intervention’s core function.
The calibration question is therefore not “Did everyone follow the script?” It is:
Did the enacted version preserve enough of the intended mechanism that the outcome can reasonably be used to judge the intervention?
School, Teacher and Student: Three Different Implementation Receipts
School
The school owns the system conditions. Were staff given enough preparation? Was implementation protected in the timetable? Were materials available? Did leadership monitor barriers and support adaptation? If the system never made delivery feasible, blaming teachers or the intervention alone is poor calibration.
Teacher or Coach
The teacher owns enacted practice. Was the intended move actually used? Was it used with adequate quality? Which adaptations were made, and why? A tick-box saying “implemented” is weaker than an observation or work sample showing what students experienced.
Student
The student is the receiver. Did the intended support reach the learner? A school can train every teacher and still fail to deliver a meaningful student experience if the intervention disappears before it reaches actual tasks, feedback or classroom routines.
Competing Explanations for a Weak Outcome
- The intervention itself is ineffective.
- The intervention is effective only for a different population or context.
- Core components were not delivered.
- Dosage was too low.
- Teachers attempted the practice but implementation quality was weak.
- Local adaptation removed an essential mechanism.
- The implementation period was too short for the intended outcome.
- The outcome measure did not capture the target effect.
- The intervention reached only part of the intended student population.
If the school does not know which of these is plausible, “it didn’t work” is not yet a diagnosis.
The Bolt Implementation-Fidelity Protocol
- Name the intended mechanism. What is the intervention supposed to change in teaching or student performance?
- Identify core components. Which parts must be preserved for the mechanism to remain plausible?
- Declare allowed adaptations. Separate sensible local tailoring from accidental dilution.
- Measure reach. Which teachers, classes and students actually received the intervention?
- Measure dosage. How much implementation occurred?
- Inspect quality. Presence is not enough; the practice must be enacted well enough to represent the design.
- Record barriers. Time, staffing, leadership, materials, turnover and training can change implementation.
- Observe outcomes at multiple levels. Teacher practice may change before student performance changes.
- Compare implementation strength with outcomes. Do stronger implementations show stronger return evidence?
- Recalibrate the verdict. Decide whether the evidence supports intervention failure, implementation failure, mixed effects, or continued uncertainty.
Worked Example: The School-Wide Feedback Routine
A school introduces a feedback routine requiring students to respond to one specific improvement instruction and resubmit a short piece of work. After a term, average writing scores do not change.
Implementation review finds that only half of departments used the resubmission step. Several teachers gave feedback but moved straight to the next task. In other classes, students corrected work but the correction was not checked.
The school cannot yet conclude that the feedback routine failed. The active element—response followed by a second performance—was absent in many classrooms.
A second cycle protects the resubmission step, simplifies the routine, monitors actual use, and collects new writing evidence. If outcomes remain flat under strong implementation, the case against the intervention becomes much stronger.
Implementation Can Also Explain Why Small Trials Shrink at Scale
Many educational interventions look strong in carefully supported trials and weaker when expanded. One reason is that scale changes the implementation environment. Coaches serve more teachers. Training becomes less intensive. Leadership attention is divided. Staff turnover appears. Contexts become more varied.
This does not mean every scale-up failure is an implementation problem. It means implementation strength is one of the competing explanations that should be measured rather than guessed.
What Current Evidence Says
The Education Endowment Foundation’s 2024 review Implementation in Education synthesises evidence on how school leaders can introduce and sustain new approaches. Its accompanying A School’s Guide to Implementation organises implementation into Explore, Prepare, Deliver and Sustain phases rather than treating implementation as a one-off launch.
Recent school-based research continues to show that implementation conditions matter. A 2024 cluster-randomised trial in Implementation Science Communications, Supporting implementation of universal prevention initiatives in K-12 schools, found that external implementation supports could improve fidelity through mechanisms including team productivity and organisational readiness.
The evidence boundary is important: fidelity is not a universal guarantee of effectiveness. A perfectly implemented weak intervention can still fail. Fidelity evidence simply lets us distinguish “the idea failed” from “the intended idea was never adequately tested.”
Common Misconceptions
- “If teachers signed the training register, implementation happened.” Attendance is preparation evidence, not classroom delivery evidence.
- “Fidelity means no adaptation.” Adaptation can be essential; the issue is whether core function survives.
- “Low fidelity means teachers are resistant.” System barriers may be decisive.
- “If outcomes improved, fidelity does not matter.” Without implementation data, it is harder to know what produced the improvement or how to reproduce it.
- “If outcomes did not improve, the programme is bad.” That conclusion requires evidence that a reasonable version of the programme was actually implemented.
For Parents: Ask What the Child Actually Received
Schools often announce new initiatives using broad labels—new feedback system, new reading programme, new AI platform, new coaching model. Parents can ask a simple calibration question: what changed in the learner’s actual day?
If the answer is unclear, the initiative may still be in an implementation stage rather than an outcome stage.
Bolt Direction Graph
Intervention design → intended mechanism → preparation → enacted practice → reach/dosage/quality → student receipt → outcome → compare fidelity with outcome → decide intervention failure, implementation failure or uncertainty → recalibrate.
Useful neighbours: Bolt Measurement Note 26 — A Coaching Conversation Is Not Yet Improved Teaching, Bolt Measurement Note 17 — One Classroom Observation Is Not the Teacher, and Bolt Measurement Note 08 — After a Very Bad Result, Improvement May Not Mean the Fix Worked.
