Wait, What? A Study Can Have Hundreds of Students and Still Have Very Little Independent Evidence
A school tests a new programme with 300 students. That sounds like a large study.
But suppose those 300 students sit inside only six classes—three intervention classes and three comparison classes.
The students are not 300 isolated experimental units. They share teachers, classmates, lessons, timetables, routines, school culture and instructional conditions. Students inside the same class tend to resemble one another more than students drawn randomly from completely different classrooms.
If the analysis pretends all 300 observations are independent, the study can look much more precise than it really is.
Quick Answer
Owned Bolt job: calibrate school and teacher intervention claims when students are clustered inside classes, teachers or schools and the number of independent assignment units is much smaller than the headline student count.
Clustering does not make school research invalid. It is normal in education. It means the analysis and sample-size interpretation must respect the level at which students share conditions. When treatment is assigned by class or school, the number of classes or schools—and the similarity of students within them—can matter more for precision than adding many extra students to the same few clusters.
The RFE question is: how much genuinely independent evidence do we have that the observed difference belongs to the intervention rather than to a small number of particular teachers, classes or schools?
Why Students in the Same Class Are Statistically Connected
Students in one classroom share a great deal:
- the same teacher;
- the same lesson sequence;
- the same classroom climate;
- the same timetable position;
- many of the same peers;
- the same school policies;
- often similar grouping or placement decisions.
Because of this, their outcomes are partly correlated. Educational researchers describe this similarity using the intraclass correlation coefficient, or ICC.
An ICC near zero means students within a cluster are not much more alike than students across clusters on the measured outcome. A larger ICC means the classroom or school membership carries more shared variation.
The important intuition is simple: twenty-five students taught by one teacher do not provide the same kind of independent evidence as twenty-five students each taught under unrelated conditions.
The Unit of Assignment Is Not Always the Unit You Measure
Student-level assignment
Individual students are assigned to conditions. If implementation can genuinely remain separate and analysis respects any remaining clustering, student-level evidence can support an individual-level treatment contrast.
Class-level assignment
Whole classes receive different conditions. The classroom becomes a critical assignment unit because every student in that class receives the same teacher-level treatment environment.
School-level assignment
Whole schools receive the intervention. Students and teachers are nested inside a much smaller number of schools, so school count becomes central to causal precision.
If a programme is assigned to six schools but analysed as though thousands of students were independently assigned, standard errors can become too small and the result can look more statistically certain than the design supports.
The What Works Clearinghouse Explicitly Adjusts for This Problem
The What Works Clearinghouse has long warned that when the unit of assignment differs from the unit of analysis, statistical tests can look more precise than they really are. Its standards explain that ignoring clustering can underestimate standard errors and overstate statistical significance.
That is not merely a research technicality. Imagine a school leader hearing that an intervention “worked for 600 students.” If those 600 students came from four intervention schools and four comparison schools, the evidence does not have the same independent breadth as 600 individually assigned students across many unrelated settings.
The students still matter. Their outcomes still matter. The correction is about how much precision and generalisability the design can legitimately claim.
Why Adding More Students to the Same Few Classes Has Diminishing Returns
In clustered designs, adding students within an existing class improves information—but usually not as much as adding genuinely new classes or schools.
The familiar design-effect relationship captures the idea:
Design effect ≈ 1 + (average cluster size − 1) × ICC
The formula is not a universal shortcut for every multilevel design, but it shows why cluster size and within-cluster similarity reduce the effective independence of a sample.
If students in the same class are highly similar, observing the thirtieth student in that class adds less new information about the intervention contrast than observing students in a completely new class taught by another teacher.
School, Teacher and Student: Three Levels of Performance Evidence
School
If only a few schools are compared, school-specific differences can dominate the estimated programme effect. Leadership, intake, timetable, staffing and other system conditions may be inseparable from treatment assignment.
Teacher or coach
If an intervention is delivered by only two teachers, a strong result may partly be a teacher effect. More students with those same two teachers do not tell us whether the method transfers to other teachers.
Student
Student-level variation still matters. But the student outcome should not be treated as though it was generated independently of the class environment that produced it.
Competing Explanations for a Strong Effect in a Six-Class Trial
- The intervention genuinely improves performance.
- One intervention teacher is unusually strong.
- One comparison class had unusual disruption.
- Baseline class composition differed.
- The programme interacts strongly with teacher quality.
- The few sampled classes happened to be atypical.
- Ignoring clustering made the estimate appear more statistically certain than it should.
A large student count cannot distinguish these explanations if the number of independent classes remains tiny.
The Bolt Clustering Calibration Protocol
- Name every level. Students sit inside classes, teachers, schools and sometimes districts.
- Name the assignment unit. Who was actually randomised or selected into the intervention?
- Inspect the number of clusters. Six schools and 3,000 students is still a six-school assignment structure.
- Estimate or justify ICC assumptions when planning. Power depends on how similar students are within clusters.
- Use analysis that respects clustering. Multilevel models, cluster-robust methods or other appropriate approaches may be needed.
- Do not report student count as though it were the whole precision story.
- Check teacher or school heterogeneity. Does the intervention work broadly or depend heavily on one cluster?
- Add clusters before merely adding more students where feasible. New classrooms or schools often provide more independent information.
- Match generalisation to sampled levels. Evidence from three teachers should not become a universal teacher claim.
- Recalibrate the RFE decision. Decide whether the study supports a promising local effect, a teacher-dependent effect, or a sufficiently replicated effect to scale.
Worked Example: 360 Students, Twelve Classes
A school network evaluates a feedback routine with 360 students across twelve classes. Six classes receive the routine and six continue usual practice.
If analysts treat all 360 students as independent, the intervention estimate may appear very precise. But every group of roughly thirty students shares one class environment. The effective information lies partly in the twelve class-level contrasts, not only the 360 student records.
The network then repeats the evaluation across thirty-six classes in several schools. The average effect remains similar and appears across many teachers.
The second result is more valuable not simply because it includes more students, but because the effect survives more independent instructional contexts.
Current Education Trials Plan Around Clustering Explicitly
Current school trials routinely design sample size around intraclass correlation. A 2025 protocol for a self-regulated learning intervention in South Australian primary schools planned for at least 56 schools and explicitly incorporated an assumed ICC into its power calculation. Another 2025 cluster-randomised school trial protocol similarly used expected school-level ICC when determining the number of participating students and schools.
IES continues to fund design work because education experiments often contain several nesting levels. Its current project on state-specific design parameters emphasises that students in the same groups tend to be more alike, that education experiments can have multiple ICCs across classrooms and schools, and that precision and minimum detectable effects depend on that structure.
The methodological lesson is mature: clustering is not optional complexity added by statisticians. It is a property of how schooling is organised.
Why This Matters for Teacher Evaluation Too
Suppose a teacher has thirty students and their average result is unusually high. Thirty students provide much more evidence than one student—but they also share one teacher and one classroom.
To infer a stable teacher effect, repeated classes and cohorts matter. If the teacher’s strong result returns across different groups, years and contexts, the teacher-level signal strengthens. One large class cannot substitute completely for repeated teacher performance.
Common Misconceptions
- “Clustered data means the students do not count.” False. Every student outcome matters; dependence changes precision, not human importance.
- “A huge student sample guarantees a powerful study.” Not when the number of independent classes or schools is very small.
- “ICC is a fixed property of a school.” It varies by outcome, age, context, covariates and level.
- “Multilevel modelling fixes a badly designed study.” Analysis can respect clustering but cannot create missing independent clusters.
- “More clusters are always easy to obtain.” School trials face real cost and recruitment constraints; the answer is honest precision, not pretending independence.
- “One teacher with many students proves the method.” It may strongly support that teacher’s implementation; transfer to other teachers remains another question.
How Do We Know?
IES’s Statistical Power Analysis in Education Research explains why field studies in education must account for multilevel clustering and how ICC, cluster size and the number of clusters affect statistical power.
The current IES project State-specific Design Parameters for Designing Better Evaluation Studies emphasises that students within classrooms and schools are correlated and that this structure directly affects precision, power and efficient sample allocation.
The What Works Clearinghouse standards explain that analysing clustered assignment as though student observations were independent can underestimate standard errors and overstate statistical significance.
A current 2025 school-trial protocol, Pragmatic clustered randomised control trial to evaluate a self-regulated learning intervention, shows how modern education trials explicitly incorporate ICC and school-level clustering into sample-size planning.
Evidence boundary: no single ICC or design-effect value applies to all school studies. Appropriate modelling depends on the assignment level, outcome, number and size of clusters, covariates and research question.
For Parents: “Hundreds of Students” Is Not the Only Number to Ask For
If an educational programme says it was tested on 1,000 students, that may be excellent evidence. It is also useful to ask: across how many teachers, classes and schools? Did the effect appear broadly, or mostly inside a few unusually strong settings?
Bolt RFE: What Should Change Next?
If clustering makes the current result fragile, the next cycle should add independent contexts rather than merely more students to the same classes. Recruit more teachers or schools, analyse at the correct levels, inspect heterogeneity, and test whether the performance effect returns when the intervention leaves the original implementation cluster.
A performance effect becomes more trustworthy when it survives not only more students, but more independent classrooms, teachers and schools.
Bolt Direction Graph
Students → classes/teachers/schools → assignment unit → within-cluster similarity → clustered analysis + precision → inspect effect across clusters → add independent contexts → observe return → recalibrate scale-up decision.
Useful neighbours include If Two Groups Start Different, the Final Difference May Not Be the Intervention, One Classroom Observation Is Not the Teacher, and A Class Score Gain Is Not Automatically the Teacher’s Effect.
