Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

Bolt Performance Calibration — Two Observers Can Watch the Same Lesson but Attend to Different Evidence

Wait, What? The Observer Can Miss the Evidence Before the Scoring Error Even Begins

Two trained observers watch the same lesson. One notices the teacher’s questioning sequence and misses several student misconceptions. The other tracks student responses closely and notices that several answers are being accepted without explanation.

By the time both observers reach the scoring rubric, they are already working from different evidence sets.

Bolt therefore separates what the lesson contained from what the observer actually attended to. Rating error can begin before the numerical rating is chosen.

Quick Answer

Owned Bolt calibration job: determine whether classroom-observation judgements are being shaped by differences in what observers visually attend to, which classroom evidence they sample, and whether that attention pattern supports an accurate teaching-quality judgement.

This is distinct from live-versus-video mode and from observation reactivity. The question here is the observer’s evidence-selection process: what entered the judgement at all?

Observation Is a Selection Problem Before It Is a Scoring Problem

A classroom contains more information than one person can process simultaneously. At any moment, an observer could attend to:

  • teacher explanation;
  • teacher movement;
  • student talk;
  • student work products;
  • one struggling learner;
  • whole-class participation;
  • questioning sequence;
  • feedback moves;
  • off-task behaviour;
  • subject-specific reasoning.

Because attention is selective, the observation process is not simply “see everything, then rate.” It is closer to: sample → interpret → score.

Observable Signs That Attention Selection May Be Distorting the Rating

  • Observers disagree most strongly on dimensions requiring attention to student thinking rather than visible teacher behaviour.
  • Post-observation discussions reveal that observers noticed different critical events.
  • One observer can justify a score with several concrete examples while another relies on a global impression.
  • Observers focus disproportionately on the teacher even when the rubric requires evidence from student responses.
  • Subject-background differences change which classroom events observers notice.
  • Immersive or wide-field observation environments change where observers look, even when the lesson itself is unchanged.

These patterns do not prove that one observer is correct and the other wrong. They show that evidence selection itself needs calibration.

Competing Explanations for Observer Disagreement

  • The observers noticed different classroom events.
  • They noticed the same events but interpreted them differently.
  • They used the rubric differently.
  • One observer had stronger subject-specific knowledge.
  • The lesson contained genuinely mixed evidence.
  • The observation environment changed what each observer could see.
  • One observer formed an early global impression and filtered later evidence through it.

The useful calibration move is to identify which stage created the disagreement: attention, interpretation or scoring.

School–Teacher–Student Triad

School

The school should train observers not only on score anchors but on what evidence to look for. If a dimension requires evidence of cognitive activation, student reasoning and response quality must enter the observation, not only teacher fluency or classroom order.

Teacher or Coach

When discussing observation feedback, the teacher can ask: “What did you see that led to this judgement?” That moves the conversation from score defence toward evidence reconstruction. If the observer did not see a critical event, the limitation should be stated rather than hidden inside certainty.

Student

Students are not passive background in an observation. Their responses often provide the strongest evidence about whether instructional moves created the intended learning opportunity. An observation system that watches only the teacher can miss the performance consequence that matters most.

The Bolt Observer-Attention Calibration Protocol

  1. Name the observation dimension. What evidence would legitimately support it?
  2. Define critical evidence zones. Teacher behaviour, student response, interaction pattern, work product or another source.
  3. Use anchor events in training. Compare what observers noticed before comparing scores.
  4. Separate seeing from interpreting. First establish whether the event entered both observers’ evidence sets.
  5. Use subject-specific calibration where needed. Some instructional evidence requires disciplinary knowledge to recognise.
  6. Compare evidence notes, not only totals. Similar scores can arise from different sampled evidence.
  7. Resample if important evidence was missed. Another lesson or video segment may be needed.
  8. Do not treat gaze as quality. Looking at the teacher or student is not inherently correct; relevance depends on the event and construct.
  9. Check whether observer attention predicts rating accuracy. Use calibrated reference events where possible.
  10. Recalibrate the judgement. State whether disagreement came from evidence selection, interpretation or scoring.

Worked Example: The Student Misconception Nobody Scores

During a Mathematics lesson, the teacher asks students to explain why two algebraic expressions are equivalent. Several students give short answers. One student gives a misconception that reveals a deeper misunderstanding.

Observer A is watching the teacher and records smooth questioning, good pacing and positive classroom climate. Observer B is watching student responses and notices that the misconception is acknowledged but not explored.

The two observers disagree on cognitive activation. The first calibration question is not “Who used the rubric better?” It is “Did both observers attend to the same evidence?”

After reviewing the video and the student response, both agree that the event is relevant. They then discuss how strongly that event should influence the dimension score. The disagreement has moved from invisible evidence selection to explicit scoring judgement, where calibration is possible.

How Do We Know?

A 2026 Learning and Instruction study, How Is Preservice Teachers’ Gaze During Classroom Observation Connected to Their Assessments of Teaching Quality?, used eye-tracking with 75 preservice teachers. Observers changed their visual focus according to classroom events, and gaze behaviour predicted the accuracy of teaching-quality ratings. The gaze–accuracy relationship was stronger in immersive 360-degree video, and subject-specific background also affected observation patterns.

The earlier study Observer Ratings of Instructional Quality: Do They Fulfill What They Promise? found that observer bias contributed materially to rating variance and that trained observers did not automatically eliminate the problem. This supports treating the observer as a measurement facet rather than a neutral recording device.

Research on observation-system validity also shows that indicators and observers contribute different amounts of error depending on the instructional dimension. That reinforces the need to ask what evidence was sampled before treating the final score as self-explanatory.

Evidence boundary: eye-tracking studies do not provide a universal “correct gaze pattern.” The 2026 study involved preservice teachers, mathematics videos and controlled observation environments. Gaze is evidence about attention selection, not a direct measure of teaching quality or observer competence.

Common Misconceptions

  • “Good observers see everything.” Human attention is selective.
  • “Looking at students is always better than looking at the teacher.” Relevant attention depends on the event and construct.
  • “If observers agree, they must have seen the same evidence.” Similar scores can arise from different evidence paths.
  • “Eye-tracking can tell us whether the rating is correct.” It can reveal attention patterns, not replace the validity argument.
  • “This is the same as live-versus-video observation.” No. Mode changes access; this page owns what the observer selects from the available evidence.

What Should Change Next?

When observers disagree, reconstruct the evidence path before averaging or moderating scores. Ask what each observer saw, what they missed, and whether the disputed dimension requires another observation or a shared anchor event.

RFE: Did the next calibrated observation improve agreement because observers sampled more relevant evidence—not merely because they learned to produce the same number?

Bolt Direction Graph

Teaching event → observer attention selection → evidence noticed → interpretation → rubric score → attention/interpretation disagreement check → shared anchor or resample → calibrated observation judgement.

Useful neighbours: Live and Video Observations Can Score the Same Teaching Differently, A Classroom Observation Rubric Can Miss the Teaching It Was Meant to Measure, and The Lesson Can Change Because Someone Is Watching.