Wait, What? A Good Idea Can Make Its Own Evaluation Harder by Spreading
A school tests a new questioning routine in half its classes. The other half continues with usual teaching so leaders can compare outcomes.
Then teachers talk. They share slides in the staffroom. A control teacher tries the questioning routine after hearing colleagues praise it. Students move between classes. Department leaders adopt pieces of the intervention informally. By the end of term, the “control” group is no longer untouched.
The intervention may be spreading because it is useful. But the comparison that was supposed to tell us whether it worked has become contaminated.
Quick Answer
Owned Bolt job: calibrate school intervention claims when teaching practices, information, materials or student exposure spill from the intervention group into the comparison group.
Contamination—also called spillover, diffusion or interference in different research traditions—can shrink the observed difference between groups because the comparison group begins receiving part of the intervention. It can also reveal a real system effect: ideas travel through schools. The correct Bolt response is not to suppress professional learning at all costs. It is to design the evaluation at the right level, measure cross-group exposure, and interpret the estimated effect as the effect of the actual contrast achieved, not the contrast imagined on the planning document.
The Counterfactual Is the Thing Being Protected
When a school asks whether an intervention worked, it is asking a counterfactual question:
What would have happened to comparable students if this intervention had not been introduced?
A control or comparison group is one way to approximate that unobserved world. If the comparison group begins receiving the same practice, the contrast weakens. A small difference between groups can then mean several different things:
- the intervention truly had little effect;
- both groups improved because the intervention spread;
- the comparison group independently adopted a similar practice;
- students or teachers carried intervention knowledge across boundaries;
- the school changed another policy affecting both groups;
- the intervention worked only through system-wide spillover rather than direct assignment.
Without exposure evidence, those explanations can be mistaken for one another.
Spillover Is Not Always a Defect
Schools are social systems. Teachers are supposed to learn from one another. Students talk. Department heads spread useful practice. A strong intervention may deliberately include knowledge-sharing.
That means spillover can be both:
- a validity threat if the purpose is to estimate a clean intervention-versus-no-intervention contrast; and
- a system outcome if the purpose is to understand how practice diffuses through a school.
Good calibration therefore starts by naming the estimand—the effect the evaluation is actually trying to estimate. Direct effect on trained teachers? Total school effect? Student exposure effect? Diffusion to peers? Different questions require different designs.
What Current Educational Research Shows
Teacher-development studies often recognise contamination explicitly. A randomised study of teacher professional development in language education used two delayed-intervention comparison groups—one in the same school and another in a different school—specifically because teachers in the same staffroom could disclose intervention content to colleagues.
A 2026 field experiment in Tanzania involving 440 teachers and more than 25,000 students found that participatory-teaching training improved mathematics outcomes and also examined peer spillovers. The researchers found limited evidence of spillover to indirectly exposed teachers despite the programme intentionally encouraging knowledge sharing. That result is important because spillover should be measured rather than assumed either large or absent.
School-trial methodology has treated contamination as a long-standing design problem. Earlier teacher-administered intervention research demonstrated diffusion effects in which control students benefited from intervention practices, while large school-based efficacy trials have monitored control classrooms specifically to detect use of intervention procedures.
The practical law is stable across decades: if units can influence one another, assignment does not guarantee isolation.
School, Teacher and Student: Three Routes of Contamination
School route
Leadership may standardise parts of an intervention across the whole school, resources may be placed on shared drives, or timetable and policy changes may affect both groups. The system itself can erase the planned contrast.
Teacher route
Teachers can share strategies formally or informally. A control teacher may observe an intervention teacher, borrow a worksheet, discuss a coaching move, or independently reconstruct the practice after hearing about it.
Student route
Students can carry materials, routines, explanations or expectations between classes and friendship groups. In some interventions, this social transmission is the mechanism. In others, it makes individual-level assignment difficult to interpret.
Competing Explanations for a Small Treatment–Control Difference
- The intervention is genuinely weak.
- The control group adopted part of the intervention.
- The intervention group implemented poorly.
- Both groups received a new school-wide practice.
- Students moved or interacted across treatment boundaries.
- The effect is delayed and the current follow-up is too early.
- The outcome measure does not capture the intended change.
- The intervention produced indirect spillovers that reduced the direct contrast.
Implementation fidelity answers what happened inside the intervention group. Contamination answers a different question: what happened to the group that was supposed to provide the counterfactual?
The Bolt Contamination Calibration Protocol
- Name the intended contrast. Intervention versus ordinary practice, intervention versus another active programme, or direct versus indirect exposure?
- Map plausible transmission routes. Staffrooms, shared drives, common meetings, mixed classes, student movement and leadership policy can all matter.
- Choose the assignment level carefully. If teachers within one school will inevitably share practice, school-level assignment may protect the contrast better than teacher-level assignment.
- Measure actual exposure. Ask comparison teachers what practices they used; inspect materials and observations where appropriate.
- Monitor both arms. Fidelity checks should not look only at whether the intervention group complied.
- Record concurrent school changes. A new whole-school initiative can become a hidden co-intervention.
- Preserve intention-to-treat analysis. Original assignment still answers an important policy question even when crossover occurs.
- Add exposure-sensitive analysis when justified. Distinguish assigned treatment from actual receipt without pretending post-randomisation exposure is itself random.
- Test for spillover explicitly where it matters. Indirect effects can be valuable outcomes rather than nuisance alone.
- Recalibrate the RFE conclusion. State whether the evidence supports direct effect, diluted contrast, system spillover, or unresolved contamination.
Worked Example: The Questioning Routine That Spread Through the Department
A mathematics department randomises six teachers: three receive coaching on a structured questioning routine and three continue as usual. After eight weeks, the coached classes outperform controls only slightly.
Observation records then reveal that two control teachers adopted the same wait-time and whole-class response routine after discussing it during department planning. The intervention teachers implemented strongly. The comparison was no longer “routine versus no routine.” It became “coached routine versus partly diffused routine.”
The small group difference should therefore not be read as proof that questioning had little educational value. The school now has two findings: direct coaching produced strong implementation, and useful practice diffused beyond the assigned teachers.
The next evaluation could randomise at department or school level, or intentionally study diffusion rather than trying to prevent it.
Why Randomisation Does Not Make Interference Impossible
Standard causal reasoning often assumes that one unit’s treatment does not alter another unit’s outcome. In real schools, that assumption can fail because people interact.
Modern causal-inference work calls this interference or spillover. Education is especially vulnerable because classrooms are embedded in departments, year groups, families and peer networks. The same connectedness that makes schools capable of organisational learning can complicate clean causal isolation.
Bolt therefore treats network structure as part of performance evidence. A school cannot assume independence merely because a spreadsheet has separate rows.
What This Does Not Mean
- Teachers should be forbidden to share good practice. No. Evaluation design should adapt to real professional learning.
- Any spillover makes a trial useless. False. It changes which effect is estimated and can often be measured.
- A small group difference means contamination occurred. Contamination is one competing explanation, not a default excuse.
- Cluster randomisation solves every problem. Spillovers can cross school or cluster boundaries too.
- Control teachers copying a useful practice is unethical. In many settings it is professionally rational; the evaluation needs to anticipate it.
- Indirect effects are always bias. They may be meaningful system outcomes if the research question includes them.
How Do We Know?
The randomised teacher-CPD study A randomized controlled trial of the effectiveness of teacher continued professional development on student language outcomes explicitly used a same-school and an external-school delayed-intervention control to address the possibility that CPD information could spread among colleagues.
The 2026 Journal of Development Economics field experiment Participatory teaching improves learning outcomes: Evidence from a field experiment in Tanzania studied 440 teachers and more than 25,000 students, finding improved mathematics outcomes from participatory teaching and limited evidence of peer spillover despite deliberate knowledge-sharing components.
The classic education study Diffusion Effects: Control Group Contamination Threats to the Validity of Teacher-Administered Interventions directly demonstrated that control students could benefit from intervention diffusion, reducing the purity of the experimental contrast.
Large school efficacy trials have also monitored contamination explicitly. The CW-FIT efficacy trial tracked use of intervention procedures in comparison classrooms precisely because teachers within the same schools could otherwise cross conditions.
Evidence boundary: spillover size is highly context-dependent. Some interventions diffuse rapidly; others barely spread. A school should measure exposure rather than applying a universal contamination assumption.
For Parents: A Comparison Group Is Not Automatically “Untreated”
When a school says a new programme performed “about the same as normal teaching,” it can be worth asking what normal teaching became during the evaluation. Did comparison teachers adopt similar practices? Did the whole school receive related training? Did students share resources across groups?
The goal is not to undermine school experiments. It is to make their conclusions match the real educational system in which teachers and students influence one another.
Bolt RFE: What Should Change Next?
If contamination is plausible, the next cycle should improve the counterfactual. Randomise at a level that fits professional interaction, measure comparison-group exposure, separate direct and indirect effects, and decide whether diffusion should be prevented, tolerated or studied as an outcome.
The aim is not laboratory purity. It is a comparison strong enough to tell the school what actually produced better performance—and whether that improvement can spread deliberately.
Bolt Direction Graph
Intervention assignment → actual teacher/student exposure → possible cross-group diffusion → observed outcomes → measure contamination/spillover → identify achieved contrast → recalibrate effect claim → redesign assignment or study diffusion → next performance cycle.
Useful neighbours include You Cannot Judge an Intervention Before You Know What Was Actually Implemented, If the Missing Students Are Different, the Intervention Effect Can Change, and A Coaching Conversation Is Not Yet Improved Teaching.
