Wait, What? A Lesson Can Be Observed Accurately and Still Be a Weak Measure of the Teacher
A senior teacher observes a colleague for forty minutes. The rubric is completed carefully. The notes are detailed. The lesson happened exactly as recorded. Yet the resulting score may still be a weak basis for a broad judgement about the teacher.
The problem is not necessarily poor observation. It is sampling. Teaching changes with the lesson goal, class, topic, stage of the unit, student response, task type, time of year and observer. One lesson is a real performance under real conditions. It is not automatically the whole teacher.
Quick Answer
Owned Bolt job: calibrate how much a single classroom observation can justify believing about broader teaching performance.
A single observation can be extremely useful for feedback about that lesson. It becomes less defensible when stretched into a high-confidence statement about stable teaching quality without enough sampling across lessons, contexts and observers. The stronger the decision, the stronger the evidence base should be.
The Same Teacher Is Not Teaching the Same Lesson Every Time
Teaching is an enacted performance. It emerges from an interaction between teacher, students, content, goals and context. A discussion lesson creates different opportunities to observe teaching than a revision lesson. A new class may respond differently from a familiar class. A lesson introducing unfamiliar material may look different from one consolidating established knowledge.
That means an observation score contains several things at once: signal about teaching, information about the particular lesson, effects of the students and context, the observation rubric, and the observer’s interpretation of that rubric.
Recent educational measurement research makes this explicit. A 2024 study on observation systems found that rater error and bias can materially affect teaching-quality scores and that the quality of an observation score depends on what the score is being interpreted as representing. A 2026 study likewise highlights the difficulty of measuring the same instructional-quality construct consistently across different lessons.
Observation for Feedback and Observation for Judgement Are Different Measurement Jobs
Formative feedback
A single lesson can provide rich material for a coaching conversation: what students did, where explanations landed, how questions were used, what opportunities for thinking appeared, and what the teacher might test next.
Summative judgement
A high-stakes claim such as “this teacher consistently performs below standard” requires a broader evidential burden. Research on classroom observation reliability has found that repeated observations and multiple observers can be necessary when decisions extend beyond one lesson. One influential study reported that modestly reliable formative feedback required several lesson visits by different observers, while reliable summative decisions required substantially more observations.
The exact number is not a universal law. It depends on the observation system, decision, context and reliability target. The durable Bolt principle is simpler: do not make a broader claim than the sample can carry.
School, Teacher and Student: Three Different Receivers of the Same Observation
School
A school must decide what an observation score will be used for before collecting it. A tool designed to trigger professional reflection may not automatically support ranking, promotion or sanction. Sampling, observer training, calibration and decision stakes should be matched.
Teacher or Coach
The teacher should treat one observation as evidence about an enacted lesson, not a verdict on professional identity. The useful coaching question is: what pattern from this lesson should we test again under another relevant condition?
Student
Students are part of the instructional context. A teacher may demonstrate strong practice with one class and struggle to make the same practice work with another. That variation should not automatically be blamed on teacher or students. It is evidence that performance is relational and context-sensitive.
Competing Explanations for a Weak Observation Score
- The observed teaching practice really is a stable weakness.
- The lesson was unusually difficult to teach because of content or task demands.
- The observed class created an atypical instructional context.
- The teacher selected a method poorly suited to that lesson but performs differently elsewhere.
- The rubric gave heavy weight to practices that were not appropriate or visible in that particular lesson.
- The observer interpreted the rubric differently from other observers.
- The observation captured an unusually strong or weak day.
Good calibration does not protect the teacher from evidence. It protects the school from pretending that one plausible explanation has already been proven.
The Bolt Classroom-Observation Calibration Protocol
- Name the decision. Is this observation for coaching, professional development, monitoring or a consequential judgement?
- Name the sampled lesson. Record class, content, lesson purpose, task type and relevant conditions.
- Separate lesson evidence from teacher-level inference. Describe what happened before describing what it means about stable capability.
- Inspect the observer channel. Was the observer trained and calibrated? What parts of the rubric require judgement?
- Sample again. Observe a different lesson, content type or class where appropriate.
- Use another observer when stakes rise. Agreement is not perfect accuracy, but multiple calibrated viewpoints can expose idiosyncratic scoring.
- Look for stable patterns. A repeated weakness across relevant contexts carries more weight than one isolated score.
- Test a teaching change. For coaching purposes, modify one practice and inspect what changes in student response and later performance.
- Recalibrate the claim. State what the accumulated evidence supports—and what remains uncertain.
Worked Example: The “Low Questioning” Lesson
An observer scores a teacher low for classroom questioning during a worked-example lesson. The first interpretation is that the teacher does not ask enough cognitively demanding questions.
A second observation takes place during a conceptual discussion lesson. The teacher now uses extensive probing questions, waits for student reasoning, and adapts follow-ups to responses. The two observations disagree.
The correct conclusion is not that one observation was useless. The first lesson still shows that questioning was limited in that context. What changes is the broader claim: the evidence no longer supports “the teacher lacks questioning capability.” A better question becomes whether the teacher selects questioning strategically across lesson types.
What This Does Not Mean
- One observation is not useless. It can be highly valuable for precise feedback about an observed lesson.
- More observations do not automatically remove bias. Systematic rater error can persist even with repeated measurement.
- Context does not excuse poor teaching. Context helps interpret performance; it does not erase responsibility.
- Every decision does not need ten observations. The amount of evidence should reflect the purpose and stakes.
- A rubric is not the teacher. It is one measurement lens with explicit and implicit assumptions.
How Do We Know?
The 2024 open-access study Signal, error, or bias? exploring the uses of scores from observation systems shows that observation-score quality depends heavily on the interpretation being made and that rater errors and biases can be substantial. A related 2024 study, Improving the Precision of Classroom Observation Scores Using a Multi-Rater and Multi-Timepoint Item Response Theory Model, examines how rater and temporal effects can distort observation scores used for consequential decisions.
The 2026 study Observing Instructional Practice: Can We Consistently Measure Teaching Quality Constructs? focuses directly on whether observational systems measure intended teaching-quality constructs consistently across different lessons. Foundational work, Once is not enough, found that reliable decisions required repeated classroom observations and that the number needed depended strongly on whether the use was formative or summative.
The evidence boundary matters: observation reliability varies by instrument, training, subject, lesson, observer and use. No fixed number of observations should be treated as universal. The defensible rule is to align sampling quality with the strength of the intended inference.
For Parents and Students: Do Not Turn One Lesson Into a Teacher Identity
A child may come home after one lesson and say, “The teacher cannot explain.” Take the signal seriously, but make the claim smaller: what was confusing in that lesson? If the pattern repeats across topics and time, the evidence becomes stronger. If it does not, one difficult lesson should not become a permanent label.
Bolt Direction Graph
Observed lesson → define purpose → record context → separate lesson evidence from teacher inference → inspect observer quality → resample across lessons/observers → identify stable pattern → test teaching change → observe student return → recalibrate.
Useful neighbours: Bolt Measurement Note 16 — Student Feedback About Teaching Is Evidence, Not a Verdict, Bolt Measurement Note 03 — When Two Good Teachers Give Different Marks, and Bolt 15 — When the Coach Is Wrong.
