Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

Bolt Measurement Note 26 — A Coaching Conversation Is Not Yet Improved Teaching

Wait, What? A Teacher Can Leave Coaching Inspired—and Teach Almost Exactly the Same Lesson Tomorrow

A coaching conversation can feel excellent. The teacher agrees with the feedback. The next steps sound clear. The coach and teacher both leave encouraged.

None of that is yet evidence that teaching changed.

Bolt separates coaching received, teacher understanding, changed instructional practice, and changed student performance. These stages can connect—but they are not interchangeable.

Quick Answer

Owned Bolt job: calibrate whether coaching has moved from conversation into changed teaching practice and then into stronger student evidence.

Teacher coaching has strong evidence as a professional-development strategy, but its effect cannot be measured by satisfaction, attendance or agreement alone. The first important receipt is whether the intended instructional practice actually changes. The next is whether students experience or perform differently in ways consistent with that change.

Four Different Coaching Outcomes

1. The teacher understood the feedback

The teacher can explain what the coach noticed and why it matters.

2. The teacher intended to change

The teacher agrees with a next move and plans to try it.

3. Teaching practice changed

The intended instructional behaviour appears in later lessons with sufficient quality and consistency.

4. Student outcomes changed

Students subsequently show stronger understanding, participation, independent performance, retention, transfer or another declared outcome.

A professional-development system can succeed at the first two and fail at the third. It can succeed at the third without yet showing the fourth. Good performance calibration makes the stage visible instead of declaring “coaching worked” too early.

Why Agreement Is Such a Weak Receipt

Teachers may genuinely agree with coaching and still encounter barriers to implementation:

  • the new practice is difficult to execute in real time;
  • the teacher lacks suitable materials or examples;
  • the class responds differently from the coaching scenario;
  • the intended change conflicts with timetable or curriculum pressure;
  • the teacher remembers the principle but not the exact move;
  • the change is attempted once and then disappears;
  • the coach’s recommendation was reasonable in theory but poorly matched to the classroom context.

This is why coaching quality must eventually be judged against enacted teaching, not only the quality of the conversation.

School, Coach and Teacher: Three Different Performance Questions

School

A school should not evaluate a coaching programme by counting sessions or collecting only satisfaction surveys. Those metrics measure participation and perception. A serious improvement system also checks whether teacher practice changes and whether the change produces useful student return evidence.

Coach

The coach needs a falsifiable next move. “Improve questioning” is too vague. “After presenting the worked example, give all students thirty seconds to formulate a reason before sampling responses, then probe one incorrect model” is observable. The next lesson can show whether the move appeared and what happened.

Teacher

The teacher needs enough agency to adapt a practice without dissolving it. Implementation is not copying a coach. It is preserving the functional purpose of the change while fitting it to the real class.

The Bolt Coaching Receipt Chain

A useful coaching cycle can be measured through a sequence of increasingly demanding receipts:

  1. Problem receipt: coach and teacher can identify the observed instructional problem precisely.
  2. Prediction receipt: they state what should change if the coaching move works.
  3. Implementation receipt: the teacher actually attempts the new practice.
  4. Quality receipt: the practice is enacted well enough to represent the intended change.
  5. Student-response receipt: students behave or respond differently in the immediate lesson.
  6. Performance receipt: later student work changes in the expected direction.
  7. Persistence receipt: the improved teaching practice survives without the coach being present.
  8. Transfer receipt: the teacher can use the underlying principle in a different lesson or class.

The chain does not need to be bureaucratic. Even a small coaching system can ask these questions informally. What matters is that improvement becomes observable.

Worked Example: “Ask More Questions”

An observer tells a teacher that students are too passive. The initial coaching recommendation is “ask more questions.” The teacher agrees and asks twenty more questions the next day.

Student understanding does not improve. Most questions are rapid factual checks answered by the same volunteers.

The coaching conversation was implemented literally but not functionally. The coach and teacher refine the target: elicit reasoning from a broader sample, give response time, and use wrong answers diagnostically.

On the next observation, more students produce visible reasoning and the teacher catches a misconception earlier. Later independent work shows fewer errors of the targeted type.

Now there is a stronger chain from coaching → changed teaching → changed student evidence.

What Research Says About Coaching

A major causal meta-analysis of 60 teacher-coaching studies found sizeable average effects on instructional practice and smaller but meaningful effects on student achievement. The important detail is that coaching affected both teaching and students; the evidence does not reduce success to how teachers felt about the coaching.

More recent evidence on professional development reinforces this mechanism. A 2025 meta-analysis of 46 experimental mathematics and science professional-development interventions found substantial average effects on teacher knowledge and instruction. Crucially, programmes that produced larger changes in classroom instruction also produced larger average student-achievement effects, while changes in teacher knowledge alone were not significantly related to student achievement.

That is almost a direct statement of the Bolt problem: professional learning becomes educational improvement when it changes enacted practice in ways that return better evidence from students.

Competing Explanations When Coaching Appears Not to Work

  • The coaching recommendation was wrong.
  • The teacher understood it but did not implement it.
  • The teacher implemented it inconsistently.
  • The practice was implemented but not with enough quality.
  • The student outcome measure was too early or too narrow.
  • The class conditions differed from those assumed in coaching.
  • The change helped some students but harmed or did nothing for others.
  • The coaching intervention needs more time or repeated cycles.

Without implementation evidence, these explanations are easily confused.

The Bolt Coaching Calibration Protocol

  1. Start from observed teaching evidence. Avoid coaching a generic trait.
  2. Name one change. The teacher should know what will look different next lesson.
  3. Make a prediction. What student evidence should change if the move works?
  4. Observe implementation. Did the intended move happen?
  5. Judge implementation quality. Presence is not enough.
  6. Observe immediate student response. Did participation, reasoning, accuracy or another intended signal change?
  7. Check later performance. A lesson can feel better without producing stronger independent work.
  8. Fade coach dependence. Can the teacher notice and correct the same performance problem independently?
  9. Transfer the principle. Does the teacher adapt it successfully to a different context?
  10. Recalibrate. Keep, modify or abandon the coaching hypothesis based on return evidence.

Common Misconceptions

  • “The teacher agreed, so the coaching worked.” Agreement is an intention-level receipt.
  • “The teacher changed, so students must improve.” Some teaching changes are ineffective or poorly matched.
  • “Student scores did not rise immediately, so coaching failed.” The measure may be too early, insensitive or unrelated to the targeted practice.
  • “More coaching hours mean more impact.” Recent PD evidence does not support a simple more-hours-is-better rule.
  • “A coach should make the teacher dependent on coaching.” The strongest endpoint is increased teacher self-observation and independent correction.

How Do We Know?

The Annenberg Institute’s summary of The Effect of Teacher Coaching on Instruction and Achievement: A Meta-Analysis of the Causal Evidence reports pooled effects across 60 causal studies of approximately 0.49 standard deviations on instruction and 0.18 standard deviations on student achievement, while also noting that effects often shrink at scale.

The 2025 EdWorkingPaper meta-analysis A Meta-Analysis of the Experimental Evidence Linking Mathematics and Science Professional Development Interventions to Teacher Knowledge, Classroom Instruction, and Student Achievement synthesised 46 experimental studies. It found stronger teacher knowledge and instructional practice under PD, and found that larger changes in instruction were associated with larger student-achievement effects.

The evidence boundary matters: coaching effects vary by subject, design, coach quality, teacher context and scale. A strong average effect does not mean every coaching programme works. The local programme still needs its own implementation and return evidence.

For Parents and Students: Better Teaching Should Eventually Become Visible in Student Work

Families may never see the coaching process. They can still ask the most important outcome question: did the teaching change create clearer explanations, stronger feedback, better task selection, better classroom response, and later more independent student performance?

Professional development matters because teaching matters. Its success should therefore eventually be answerable to teaching and learning evidence—not attendance certificates.

Bolt Direction Graph

Observed teaching problem → coaching hypothesis → teacher understanding → implementation → implementation quality → student response → later performance → coach fade → transfer to new context → recalibrate.

Useful neighbours: Bolt Measurement Note 17 — One Classroom Observation Is Not the Teacher, Bolt Measurement Note 16 — Student Feedback About Teaching Is Evidence, Not a Verdict, and Bolt 32 — The Teacher’s Final Job Is to Become Less Necessary.