Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

Bolt Measurement Note 25 — A Class Score Gain Is Not Automatically the Teacher’s Effect

Wait, What? The Class Improved. That Does Not Mean We Have Measured the Teacher Yet

A class begins the year with an average of 58 and ends with an average of 72. The teacher may have done excellent work. The students may have learned a great deal. But the 14-point gain is not automatically a 14-point measure of the teacher.

Student performance changes for many reasons at once: prior attainment, curriculum exposure, tutoring, attendance, peer effects, family support, motivation, maturation, assessment design, class composition, school systems, and teaching. If the purpose is to estimate the teacher’s contribution, Bolt must separate student growth from teacher effect.

Quick Answer

Owned Bolt job: calibrate what class score growth can and cannot justify believing about a teacher’s contribution to student performance.

Class-level score gains are valuable evidence, but they are not a direct causal measurement of the teacher. Better attribution requires attention to starting points, student assignment, classroom composition, test properties, other supports, repeated years where possible, and complementary evidence about teaching practice. The stronger the personnel or school decision, the more dangerous it is to treat a raw gain as the teacher.

Three Different Quantities That Are Often Collapsed

1. Student performance

What students actually produced on an assessment under declared conditions.

2. Student growth

How performance changed between two or more measurement points.

3. Teacher contribution

The portion of that change that can defensibly be attributed to the teacher rather than to other influences.

The first is observable. The second is a comparison. The third is an attribution problem.

That distinction is why educational researchers developed value-added models rather than simply ranking teachers by end-of-year averages. Modern value-added work attempts to account for prior performance and other contextual factors so that teacher contribution is estimated more fairly. Even then, the estimate remains a model-based quantity with uncertainty rather than a direct reading from the classroom.

Why Raw Gains Can Mislead

  • Students do not start at the same point. A class with major prerequisite gaps has a different performance landscape from a class already near the assessment ceiling.
  • Students are not randomly assigned to teachers. Ability grouping, subject combinations, behaviour needs, timetable structures and school decisions can affect who enters a class.
  • Other teaching exists. Students learn from previous teachers, tutors, peers, parents, textbooks, online materials and school-wide programmes.
  • The test may capture only part of the teacher’s contribution. Attendance, persistence, subject interest, discussion quality and other outcomes may matter without appearing in one test score.
  • Year-to-year estimates can be noisy. Class size, unusual events and measurement error can move a teacher’s apparent effect.

A Better Teacher-Effect Question

Instead of asking, “How many marks did this teacher add?” ask:

Given where these students started, who they were, what they experienced, and what was measured, how much evidence supports the conclusion that this teacher contributed to the observed improvement?

This framing is slower but more honest. It allows strong teachers to be recognised without pretending that every influence on student performance belongs to them.

School, Teacher and Student: The Same Gain Means Different Things

School

A school deciding whether a teaching approach or teacher is effective should use multiple measures and preserve uncertainty. If student growth is included, the school should inspect prior attainment, class composition, score precision, assessment comparability and repeated evidence rather than turning one cohort’s raw gain into a permanent teacher ranking.

Teacher or Coach

A teacher should treat class growth as a return signal, not a self-congratulating number or an accusation. Which students improved? On what tasks? Under what support? Which parts of teaching changed? Did the same instructional moves produce useful evidence across different classes or years?

Student

Student growth belongs partly to the student. A strong teacher can create better conditions, instruction and feedback; the learner still performs the learning and the later task. Bolt avoids a model in which every success belongs to the teacher and every failure belongs to the child.

Competing Explanations for a Large Class Gain

  • The teacher improved student learning substantially.
  • The class began unusually low because the baseline measure under-estimated capability.
  • The second assessment was easier or sampled different content.
  • Students received substantial external support.
  • The class composition changed during the year.
  • Students became more familiar with the test format.
  • A school-wide intervention improved performance across several classes.
  • The teacher genuinely improved some outcomes that the test captures and others that it does not.

These are not reasons to distrust improvement. They are reasons to avoid causal overclaiming.

The Bolt Teacher-Contribution Calibration Protocol

  1. Define the outcome. Which student performance is being used to represent teacher impact?
  2. Check starting points. Compare students using relevant prior performance, not only end scores.
  3. Inspect assignment. Were students sorted into classes in ways likely to affect growth?
  4. Check the assessment. Are baseline and follow-up scores comparable and sufficiently reliable?
  5. Map other influences. Tutoring, school programmes, attendance and major support changes may matter.
  6. Estimate uncertainty. A teacher-effect estimate should not be reported with more precision than the data support.
  7. Repeat across cohorts where possible. Stability across years strengthens inference.
  8. Triangulate with teaching evidence. Observation, student feedback and work samples can help explain how improvement may have occurred.
  9. Look beyond one outcome. Teachers may affect grades, attendance, behaviour or other student outcomes differently from test scores.
  10. Match use to evidence. Low-stakes coaching can tolerate more uncertainty than dismissal, promotion or public ranking.

Worked Example: Two Teachers, Two Very Different Classes

Teacher A’s class rises from 80 to 86. Teacher B’s class rises from 50 to 65. A raw-gain ranking declares Teacher B superior because the gain is fifteen points instead of six.

But Teacher A’s students began near the top of a test with limited room to show further growth. Teacher B’s class had substantial unmastered material and also received a new after-school support programme. The assessments differ slightly in content emphasis.

The raw gains remain real. The ranking does not. We need a better model before attributing the difference to teachers.

If Teacher B’s students repeatedly outperform reasonable predictions across cohorts and the pattern aligns with observed instructional strengths, the teacher-effect claim becomes stronger. If Teacher A’s students show excellent later outcomes not captured by the original test, the evidence may also broaden our understanding of Teacher A’s contribution.

Why Value-Added Helps—and Why It Is Still Not the Whole Teacher

Value-added models were designed to improve on crude comparisons by controlling for prior achievement and, depending on the model, other student and school characteristics. Large causal studies have found that carefully specified value-added measures can contain real signal about teachers’ contributions to test-score growth.

At the same time, current research continues to investigate what value-added misses. A 2024 Institute of Education Sciences project on lasting teacher impacts explicitly notes concern that test-based value-added does not capture teachers’ full contributions and is developing measures that combine test and non-test outcomes such as attendance, discipline, grades and promotion.

This is an important Bolt principle: a useful measure does not need to be a complete measure.

Common Misconceptions

  • “Student growth has nothing to do with the teacher.” False. Teacher effects on achievement are well supported; attribution simply requires careful modelling.
  • “Value-added gives the true teacher effect.” It gives an estimate based on assumptions, data and a defined outcome.
  • “Multiple measures solve everything.” Multiple weak measures do not automatically create one strong judgement.
  • “If a teacher’s score changes by year, the measure is useless.” Some real teacher performance may change, and some variation reflects noise; the job is to estimate both.
  • “The teacher owns all class outcomes.” Student performance is produced by a system of influences, including the student.

How Do We Know?

The Institute of Education Sciences currently supports the 2024–2028 project Understanding Lasting Teacher Impacts, which is developing and validating measures that combine test and non-test outcomes while accounting for classroom composition and student sorting.

Foundational causal evidence from the American Economic Association includes Measuring the Impacts of Teachers I: Evaluating Bias in Teacher Value-Added Estimates, which found little bias in appropriately specified value-added forecasts in the setting studied. IES research also continues to examine threats from non-random student assignment in Value-Added Models and the Measurement of Teacher Quality.

IES work on Reducing Bias and Improving Efficiency of Estimated Teacher Effects from Value-Added Models identifies non-random assignment, assumptions about persistence and test scaling as continuing methodological concerns. Earlier IES work on error rates also shows why single-year teacher and school estimates can be imprecise.

The evidence boundary matters: value-added evidence is strongest for the outcomes and contexts actually modelled. It should not be stretched into a universal measure of teacher quality, character or worth.

For Parents: “The Class Improved” Is Good News—Then Ask What We Can Attribute

Parents do not need a statistical model to ask better questions. If a class improves strongly, celebrate it. Then ask what else changed: teaching, assessment, support, attendance, composition, tutoring, curriculum or student effort. The aim is not to dilute credit. It is to learn which conditions should be preserved for the next cohort.

Bolt Direction Graph

Student outcomes → observed growth → inspect starting points and assignment → check assessment and other supports → estimate teacher contribution with uncertainty → triangulate with teaching evidence → repeat across cohorts → recalibrate teacher-performance claim.

Useful neighbours: Bolt Measurement Note 17 — One Classroom Observation Is Not the Teacher, Bolt Measurement Note 16 — Student Feedback About Teaching Is Evidence, Not a Verdict, and Bolt Measurement Note 02 — Before You Call It Improvement, Check Whether the Scores Are Comparable.