Bolt Series · Human Performance Calibration · Article 36
We have spent 35 articles asking whether the learner is calibrated. What about the school?
A student can misunderstand their own capability.
A parent can misunderstand it.
A coach can misunderstand it.
A teacher can misunderstand it.
So can a school.
Not because schools are careless.
Because a school is also an observation system.
It collects scores.
Teacher judgements.
Classwork.
Behaviour.
Past records.
And from those signals it builds a model of the learner.
That model can be excellent.
It can also be incomplete, outdated or wrong.
First, a correction: teachers are not generally bad judges
It would be easy to turn this article into a fashionable argument that schools do not understand children.
The evidence does not justify that.
A 2024 psychometric meta-analysis re-examining teacher judgement accuracy concluded that earlier reviews had likely underestimated how accurately teachers judge students’ academic achievement and overestimated how much accuracy varies across studies.
Teachers see students repeatedly.
They observe work across time.
They often have far more information than one test provides.
So the correct starting point is not:
“Teacher judgement is unreliable.”
It is:
Teacher judgement is valuable evidence — and, like every evidence source in Bolt, it must remain correctable.
Accuracy and error can coexist
A system can be accurate most of the time and still make consequential mistakes.
Airports are generally excellent at routing luggage.
Some bags still arrive in the wrong city.
Medical tests can be highly useful.
False positives and false negatives still matter.
The same logic applies in education.
A teacher can usually judge a class well and still underestimate one particular learner.
A school can use generally valid assessments and still make an over-large inference from one result.
The question is not whether error exists.
It is whether the system can detect and repair it before the error becomes a pathway.
An external model can affect the opportunities that generate future evidence
This is where school miscalibration becomes more serious than an ordinary mistaken estimate.
Suppose a teacher believes a student is weak.
The student receives easier work.
Fewer difficult questions.
Less opportunity to demonstrate advanced reasoning.
The original estimate now has fewer chances to be contradicted.
APA guidance on teaching and learning explicitly notes that teacher expectations can influence students’ opportunities to learn, motivation and outcomes. It also cautions that inaccurate expectations can sometimes become self-fulfilling through differential treatment.
That does not mean every low expectation creates a large self-fulfilling prophecy.
A major review of the teacher-expectation literature concluded that self-fulfilling effects do occur but are typically small and often dissipate rather than accumulating without limit.
Still, even a small average effect can matter for particular students and decisions.
Bias can be domain-specific
One of the most useful findings for Bolt is that judgement bias does not always operate as one global prejudice.
A 2023 cross-national study using longitudinal data from England, Germany and the United States found domain-specific gender bias in teacher judgements: a positive bias for girls in language and for boys in Mathematics. Those biases partly mediated later gender achievement gaps.
This is exactly the kind of result calibration theory should expect.
The external observer is not simply “biased” or “unbiased.”
Accuracy can vary by domain, context and student characteristic.
That means the school’s model needs the same resolution we have demanded from the learner’s self-model.
Student characteristics can alter judgement accuracy too
A 2025 study examined teacher and parent judgements of children’s cognitive abilities in relation to characteristics including special educational needs, giftedness and socioeconomic status.
The reason this matters is not that a demographic characteristic automatically produces a wrong judgement.
It is that characteristics unrelated to the capability we are trying to estimate can sometimes influence the estimate.
When that happens, the observation system is no longer reading only the learner.
It is also reading its own expectations.
A strong school should be calibrated not only about students, but about the reliability and possible biases of its own judgement process.
The test can be right and the school inference can still be wrong
Imagine a student scores poorly on a valid, correctly marked examination.
The score may be completely accurate as a record of that performance.
The school then concludes:
“This student is not capable of higher-level work.”
That is a larger claim.
Modern educational measurement insists that score interpretation and use require evidence appropriate to the claim being made.
The more ambitious the inference, the more evidence is required.
A school can therefore be miscalibrated even when the test itself is functioning properly.
The error occurs in the jump from:
“This performance was weak.”
to:
“This defines the learner’s wider capability or future potential.”
The most dangerous error is a model that stops generating opportunities to be disproved
Suppose a school believes a student cannot handle advanced work.
If the student is never given advanced work again, what evidence could prove the model wrong?
This is a deeper measurement problem.
The model has begun controlling the evidence environment.
A well-calibrated institution therefore needs periodic opportunities for models to fail.
- Fresh tasks that are not merely repetitions of old placement assumptions.
- Opportunities to attempt harder work where appropriate.
- Multiple observers when an important judgement remains uncertain.
- Evidence from different task formats and conditions.
- Reassessment after substantial learning or development.
- A route for the learner to challenge an interpretation with new performance evidence.
The point is not to abolish professional judgement.
It is to keep professional judgement empirical.
What happens when the student and school disagree?
The student says:
“I can do more than this.”
The school says:
“Our evidence says otherwise.”
Do not solve the disagreement by choosing the student because self-belief is empowering.
Do not solve it by choosing the institution because authority is experienced.
Ask what each model predicts.
Then create a fair task capable of distinguishing the predictions.
If the student succeeds repeatedly, the school model should move.
If the student does not, the student’s self-model should move.
If the result is mixed, both models may need more resolution.
No observer earns the right to stop updating.
A school should know how uncertain it is
Schools often have to make decisions before perfect evidence exists.
That is unavoidable.
But there is a large difference between:
“We currently think this learner needs more foundational support, and our confidence is moderate because the evidence is mixed.”
and:
“This is a weak student.”
The first statement preserves uncertainty and future correction.
The second converts a provisional model into an identity.
Institutions need calibrated uncertainty just as learners do.
The school can also underestimate improvement
A student changes.
But school records are historical.
A teacher reads last year’s report.
A placement decision came from an earlier phase of development.
The learner is still being interpreted through yesterday’s evidence.
This is the institutional version of the underconfident learner who continues to say “I’m bad at fractions” after months of strong new performance.
A school model can become stale too.
What would a calibrated school do?
Not assume every teacher is biased.
Not assume every score is incomplete to the point of uselessness.
Not accept every learner self-report uncritically.
Instead:
- Use teacher judgement as valuable evidence.
- Use well-designed assessments for the claims they can support.
- Separate observed performance from broader capability inference.
- Notice when demographic or contextual information may be contaminating judgement.
- Keep important decisions open to new evidence.
- Give learners fair opportunities to demonstrate change.
- Require stronger evidence for more consequential decisions.
- Track whether institutional predictions actually come true.
That last item is critical.
If a school predicts that a learner cannot succeed at a certain level and the learner later succeeds, the school should learn from the prediction error.
Calibration is not only something we demand from children.
Adults and institutions have to close the loop too.
Why this matters for education
A student’s wrong self-estimate can waste effort.
A school’s wrong estimate can alter opportunity.
That makes external calibration a serious responsibility.
The answer is not to distrust schools.
It is to design school judgement so that trust is deserved.
A good school does not prove its judgement by never being wrong. It proves its judgement by measuring carefully, knowing the limits of its evidence, and changing its mind when the learner produces a stronger receipt.
That is the institutional version of Bolt.
The school should know the learner.
And it should keep checking whether it knows the learner as well as it thinks it does.
Evidence and further reading
- Meta-Analysis — Teachers’ Judgment Accuracy
- APA — Teachers’ Expectations and Students’ Opportunities to Learn
- Review — Teacher Expectations and Self-Fulfilling Prophecies
- Teacher Judgements and Gender Achievement Gaps Across Three Countries
- Special Educational Needs, Socioeconomic Status and Judgements of Cognitive Ability
- Review — The Impact of Teacher Expectations
- NCME — Validity and Educational Testing
