Bolt Series · Human Performance Calibration · Article 29
What if every observer tells a different story?
A student says:
“I understand this topic. I just had a bad paper.”
The teacher says:
“I think the understanding is weaker than you realise.”
The test says:
52%.
The last three homework sets say:
85%, 88%, 91%.
A second teacher says:
“The student explains the concepts very well in class.”
Now what?
This is where calibration becomes genuinely difficult.
Not when the evidence agrees.
When it does not.
Do not solve disagreement by choosing your favourite observer
There are several tempting shortcuts.
Trust the student because they know themselves best.
Trust the teacher because the teacher is the expert.
Trust the test because the test is objective.
Trust the average because more numbers must be better.
Each shortcut can fail.
The student has privileged access to internal experience but can misread it.
The teacher has experience and an outside view but can carry an outdated or mistaken model.
The test provides a powerful external receipt but only for what that test actually measured under those conditions.
An average can stabilise noise and still hide a local failure that matters.
Conflicting evidence is not a nuisance to remove. It is a signal that the current model is too simple.
First ask whether the evidence is actually about the same thing
Suppose homework is excellent and the examination is weak.
That looks contradictory.
But perhaps the homework allows notes, examples, unlimited time and repeated teacher checking.
The examination requires unaided retrieval, transfer and pacing.
Now the two measurements may not be contradicting each other at all.
They may be measuring different performance conditions.
The apparent conflict becomes:
“This learner performs strongly with support and weakly when retrieval, transfer and time pressure are combined.”
That is not contradiction.
That is diagnosis.
Then ask how trustworthy each source is for this particular question
Evidence should not be counted as though every source has equal diagnostic power.
Research on belief updating shows that people can and do weight information partly according to source reliability. A 2025 series of experiments found that participants updated beliefs differently depending on the credibility of conflicting sources and appropriately discounted sources whose information had been corrected.
That principle transfers carefully to learning.
If the question is:
“Can the student solve this class of problems independently under time pressure?”
then an unaided timed task is more diagnostic than a feeling of familiarity.
If the question is:
“Why did the student abandon a correct method halfway through?”
then the learner’s own report may contain information the final score cannot reveal.
Evidence should therefore be weighted by its relationship to the claim being tested.
Confidence changes how people process disagreement
There is another reason conflicting evidence is hard.
People do not receive contradictory feedback neutrally.
Research on probabilistic learning has found that higher confidence can reduce how strongly feedback is processed and how much beliefs are updated afterwards.
Another recent study found that newly formed self-beliefs can become resistant to contradictory feedback, especially as confidence in those beliefs grows.
This matters in both directions.
The overconfident learner can say:
“The test is wrong. I know I’m good at this.”
The underconfident learner can say:
“Those three good results do not count. I know I’m bad at this.”
Both protect the existing self-model by rejecting inconvenient receipts.
Healthy calibration requires something harder:
Allow evidence to challenge the model without allowing one surprising observation to seize control of the whole model.
Do not vote. Design a discriminating next test.
Imagine two explanations for the 52% examination.
Explanation A: the learner understands the concepts but cannot retrieve and execute them efficiently under time pressure.
Explanation B: the learner’s conceptual understanding is weaker than homework performance suggests because support has been masking gaps.
Which explanation wins?
Do not ask who argues more confidently.
Design a task that makes the explanations predict different outcomes.
For example:
- Give an unfamiliar problem with no time pressure.
- Ask for a verbal explanation of the underlying concept.
- Remove notes and worked examples.
- Then add time pressure on a second equivalent task.
If conceptual explanation and untimed transfer are strong but timed execution collapses, Explanation A gains weight.
If understanding breaks even without time pressure and support, Explanation B gains weight.
The disagreement has created an experiment.
Sometimes the evidence should remain unresolved for a while
Education often feels pressure to label quickly.
Strong.
Weak.
Ready.
Not ready.
But sometimes the most accurate state is:
“We do not yet know which explanation is better.”
That is not indecision if we know what evidence would resolve it.
It is calibrated uncertainty.
The student should learn how to hold disagreement too
A mature learner should be able to say:
“My last test says one thing, my practice history says another, and I am not yet sure why. I need another task that separates knowledge from time pressure.”
That is a very sophisticated sentence.
It avoids three common errors:
- defending the preferred self-image;
- surrendering to the latest score;
- pretending uncertainty does not exist.
The learner remains correctable without becoming unstable.
Why this matters for education
The world does not always give us one clean measurement.
Teachers disagree.
Tests disagree.
Past and present performance disagree.
Feelings and outcomes disagree.
The educational goal cannot be to remove all disagreement.
It should be to teach people how to reason through it.
When evidence disagrees, do not ask which observer owns the truth. Ask which claim each source can support, how reliable that source is for the claim, and what next receipt would force the competing explanations apart.
That is not merely good test interpretation.
It is a way of keeping our representation of a learner correctable by the world.
Evidence and further reading
- Belief updating in the face of misinformation: the role of source reliability
- Confidence regulates feedback processing during human probabilistic learning
- Initial expectations and confidence affect the formation and revision of self-beliefs
- Thinking about believing: counterevidence and belief updating
- NCME — Validity and Educational Testing
