Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

Bolt Performance Calibration — The Lesson Can Change Because Someone Is Watching

Wait, What? The Observation Can Become Part of the Lesson

A school leader enters a classroom with a rubric. The teacher knows the lesson matters. Students notice the visitor. Instructions become more deliberate. Transitions tighten. The teacher may choose safer examples, check behaviour more carefully, or avoid an experiment that could fail.

The observer may record the lesson accurately. Yet the lesson may no longer be fully representative of the teacher’s ordinary teaching.

This is the observation-reactivity problem: measurement can sometimes alter the performance being measured.

Quick Answer

Owned Bolt job: calibrate teaching-performance evidence when the presence, visibility, purpose or stakes of an observation may change teacher or student behaviour.

Observer reactivity should neither be assumed nor ignored. Empirical classroom research is mixed: some studies find little detectable change in measured teaching quality under observation, while other work suggests that observation mode, evaluation stakes, emotions and professional risk can alter behaviour. The correct Bolt response is therefore conditional: make the observation conditions explicit, sample more than once, compare with routine evidence, and reduce the temptation to treat one observed lesson as untouched reality.

Observation Has Two Jobs That Should Not Be Confused

Observation as evidence

The observer wants to know what happened: how explanations were structured, which students participated, how feedback operated, whether misconceptions surfaced, and how the teacher responded.

Observation as intervention

The act of observing can itself change conditions. A teacher may prepare differently, students may behave differently, or the classroom may temporarily become more controlled. When the observation is high-stakes, the incentive to present an unusually polished lesson can increase.

Those two jobs can coexist. The lesson is still real. The teacher really did produce that performance. The calibration question is whether that performance should be generalised to ordinary practice.

Why the Stakes Matter

A peer visit for collaborative learning does not create the same behavioural incentives as an observation tied to promotion, appraisal, public ranking or employment consequences. The more consequential the evaluation, the more plausible it becomes that teachers will optimise for the observation system itself.

A 2026 review of teacher-evaluation research described a continuing tension between professional-growth and accountability purposes. The review notes evidence that high-stakes evaluation can reshape teacher practice and that observed practice may differ from more ordinary instruction when the evaluation period passes.

This does not make accountability illegitimate. It means the measurement design should acknowledge that incentives can affect the sample being observed.

The Evidence Is More Nuanced Than “People Behave Differently When Watched”

The popular story is simple: observation always changes behaviour. Classroom evidence is not that clean.

A study involving 447 students in 24 classes compared video-recorded and non-recorded lessons and found no significant differences in reported teaching quality or teaching practices, although teacher and student emotions, cognition and behaviour showed some differences. Eye-tracking evidence suggested that some initial reactivity declined quickly.

An earlier secondary-school case study proposed remote live-video observation as one way to reduce visible observer effects. Other work on treatment integrity found improved teacher implementation after performance feedback but no clear difference between observer-present and observer-absent conditions.

The world-class conclusion is therefore not “observation is contaminated.” It is: reactivity is a plausible measurement condition whose size depends on purpose, visibility, context and stakes.

School, Teacher and Student: Three Different Parts of the Reactivity Field

School

The school controls the stakes, observation frequency, notice period, rubric, observer training and whether observation is primarily developmental or evaluative. If a system creates unusually performative lessons, that is partly a system-design problem rather than proof of teacher dishonesty.

Teacher or Coach

The teacher should be able to distinguish “I can teach like this once” from “this is now routine practice.” A coach should look for practices that survive across ordinary lessons, not merely the most polished observation.

Student

Students can also respond to observation. They may become quieter, more cooperative, more anxious, more attentive, or more performative. Student behaviour observed during a formal visit should therefore not automatically be treated as the normal classroom baseline.

Competing Explanations for an Unusually Strong Observed Lesson

  • The teacher genuinely teaches at this level routinely.
  • The observer happened to see an unusually strong lesson.
  • The teacher prepared more heavily because the observation was known in advance.
  • The students changed their behaviour because an observer was present.
  • The teacher selected a lesson type that displays rubric-friendly practices unusually well.
  • The observer focused on visible strengths and missed less visible weaknesses.
  • The observation itself helped the teacher concentrate and perform better.

No single explanation should be chosen from suspicion or loyalty. Resampling under different conditions is the discriminating move.

The Bolt Observation-Reactivity Calibration Protocol

  1. Name the purpose. Coaching, research, appraisal and accountability create different incentives.
  2. Declare the observation conditions. Was the visit announced? Was a camera present? Who knew the stakes?
  3. Record contextual changes. Did class size, seating, lesson type, student grouping or resources differ from routine practice?
  4. Use repeated observations. One performance is too vulnerable to exceptional preparation or an exceptional day.
  5. Vary the sample. Observe different classes, topics, lesson phases or days where appropriate.
  6. Compare with routine evidence. Student work, planning, feedback records and ordinary classroom artefacts can test whether the observed practice persists.
  7. Reduce unnecessary performativity. Normalise observation, separate developmental visits from punitive uses where possible, and avoid rewarding theatrical compliance with a rubric.
  8. Use less intrusive observation when the construct allows it. Remote or video methods may reduce some forms of reactivity, although they introduce their own measurement limits.
  9. Ask what survives after the observer leaves. If the improvement is real, later ordinary lessons should continue to show it.
  10. Recalibrate the claim. State whether evidence supports “strong observed lesson,” “stable teaching practice,” or something in between.

Worked Example: The Perfect Observed Discussion

During a formal observation, a teacher runs an excellent discussion. Students explain reasoning, challenge one another and build on prior answers. The observer scores classroom dialogue very highly.

Two weeks later, a low-stakes peer visit finds that ordinary lessons rely heavily on teacher explanation and rapid volunteer responses. Student work also shows few traces of the reasoning routines seen during the formal observation.

The first lesson was not fake. It demonstrated capability: the teacher can facilitate strong dialogue. The second sample changes the interpretation: the practice is not yet reliably embedded.

The coaching target becomes clearer—move from peak observed performance to stable routine performance.

Peak Performance and Typical Performance Are Both Useful

A highly prepared observation can answer an important question: What can this teacher do when preparation and attention are concentrated?

An ordinary lesson answers a different question: What does this teacher usually sustain under normal workload?

School improvement needs both. Peak performance can reveal the next attainable standard. Typical performance reveals what students actually receive repeatedly.

Common Misconceptions

  • “Observed lessons are fake.” Too strong. They are performances under observation conditions.
  • “Unannounced observations reveal the truth.” They reduce preparation effects but still sample one lesson and can create other problems.
  • “Video solves reactivity.” Cameras can reduce the presence of an in-room observer but may still affect behaviour and change what evidence is visible.
  • “If reactivity is small on average, it never matters.” High-stakes or unusual settings may behave differently.
  • “If a teacher performs better when watched, the gain is worthless.” It may reveal genuine capability that coaching can help stabilise.

How Do We Know?

The classroom study Reactivity effects in video-based classroom research compared recorded and non-recorded lessons across 24 classes. It found no significant differences in reported teaching quality or teaching practices, but did find differences in emotions, cognition and behaviour, with some initial reactivity appearing to diminish rapidly.

The study Live video classroom observation: an effective approach to reducing reactivity in collecting observational information for teacher professional development examines remote observation as a way to reduce visible observer effects. Earlier applied work, Using Performance Feedback to Improve Treatment Integrity of Classwide Behavior Plans, found no clear observer-present versus observer-absent difference in teacher implementation in the setting studied.

For the high-stakes side of the problem, the 2026 review The Enduring Tension Between the Professional Growth and Accountability Purposes of Teacher Evaluation synthesises research showing that consequential evaluation can reshape professional practice and create a tension between authentic routine teaching and observed evaluative performance.

The evidence boundary matters: classroom reactivity is not a universal constant and should not be assigned a fixed correction factor. The safest inference is conditional and empirical: check whether observed practices survive across settings, time and reduced observation salience.

For Parents: One Open Lesson Is a Window, Not the Whole House

Parents may attend open classrooms or view selected teaching clips. These can be informative, but they are curated or observed moments. The strongest questions are about routine: How often does this happen? What do students produce over time? Does feedback return? Does independent performance improve?

Bolt Direction Graph

Teaching performance → observation conditions → possible reactivity → observed evidence → resample under different visibility/stakes → compare routine artefacts → test persistence → distinguish peak from typical practice → recalibrate teacher-performance claim.

Useful neighbours: One Classroom Observation Is Not the Teacher, Live and Video Observations Can Score the Same Teaching Differently, and A Coaching Conversation Is Not Yet Improved Teaching.