Bolt Series · Human Performance Calibration · Article 26
A surprising score should change the model — but by how much?
A student has scored between 78% and 84% across several assessments.
Then one paper comes back:
54%.
What should we do?
One response is to ignore it.
“That is not really you.”
Another response is to let the new result rewrite everything.
“You are actually a 54% student.”
Both reactions are too simple.
The 54 is real evidence.
But evidence has to be weighted.
A score is an observation, not a total replacement for history
Educational measurement has always had to deal with a basic problem:
Observed scores are not perfectly stable.
Tests sample questions.
Performance varies.
Measurement contains error.
Different tasks expose different parts of a domain.
Modern validity theory therefore asks whether evidence supports a particular interpretation and use of a score, rather than treating the number as carrying unlimited meaning by itself.
The latest edition of Educational Measurement makes this distinction explicit: generalisation from an observed score depends on the sample of questions or observations, and errors of measurement affect how confidently we can move from that sample toward broader conclusions.
So one result should enter the model.
It should not automatically become the model.
The right question is not “Which score is the real one?”
Suppose the student’s recent history is:
- 81%
- 79%
- 84%
- 80%
- 54%
Which score is real?
All of them.
They are real performances under real conditions.
The more useful question is:
What model of the learner best explains the whole pattern?
Perhaps the 54 exposed a genuine weakness that the earlier papers did not sample.
Perhaps this paper demanded much more transfer.
Perhaps time pressure was unusually high.
Perhaps the learner had an abnormal state that day.
Perhaps the earlier scores were inflated by heavy support.
The new result raises hypotheses.
Then we test them.
Surprising evidence deserves attention, not panic
A surprising result is often especially valuable.
It tells us the current model failed to predict something.
If a student reliably performs at 80 and suddenly produces 54, the model now has an unexplained residual.
That is worth investigating.
But investigation is different from identity revision.
A surprising observation should increase curiosity before it increases certainty.
Ask how diagnostic the result is
Not all evidence deserves equal weight.
A result becomes more informative when:
- the task closely matches the capability we are trying to infer;
- the assessment is reliable enough for the decision being made;
- the work was completed under known conditions;
- the score is supported by visible response patterns rather than only a total;
- similar weaknesses appear again on another task;
- the learner’s own report and external evidence help explain the mechanism.
A single highly diagnostic task can sometimes teach us a great deal.
A single noisy or poorly matched task may teach us much less.
This is why “one result is never enough” would also be too crude.
The correct weight depends on the quality of the evidence.
Repeated evidence reduces the chance that we are chasing noise
Reliability research exists partly because scores can vary under repeated measurement.
Low reliability means large changes can appear on retesting even when the underlying thing we care about has not changed very much.
In performance assessment, generalizability theory goes further by asking how different sources of variation—tasks, raters, occasions and their interactions—contribute to an observed result.
That is a sophisticated version of a common-sense educational principle:
Before making a large judgement about a person, find out whether the pattern survives another observation.
But do not average away a real weakness
There is an opposite danger.
A learner has excellent overall results.
One new assessment reveals a serious recurring failure in one important subskill.
The adult says:
“Ignore it. Your average is still high.”
That can be just as damaging as overreacting.
An overall average can hide a local weakness.
If the new task was specifically designed to test transfer, and transfer failed cleanly, the observation may deserve substantial weight even if the total history is strong.
The question is never merely:
How many pieces of evidence do I have?
It is also:
What does each piece actually diagnose?
The learner should update this way too
Students often let one event rewrite their self-model.
One difficult paper:
“I’m terrible at this subject.”
One excellent paper:
“I’ve mastered everything.”
Both are unstable ways to learn about yourself.
A stronger response sounds like this:
“This result is different from what I expected. I am going to work out what it should change in my current model, and what it should not change yet.”
That is intellectual discipline.
A practical update rule
When a surprising result appears, ask:
- What happened? Describe the performance before explaining it.
- How different is this from the previous pattern?
- What exactly did this task measure?
- What mechanisms are visible in the work?
- What conditions were different?
- What new prediction does each explanation make?
- What next task would distinguish those explanations?
Then collect the next receipt.
This is better than either defending the old model or surrendering to the newest number.
Why this matters for parents and teachers
Important educational decisions can alter opportunity.
Difficulty level.
Subject choices.
Support.
Expectations.
One observation should therefore be interpreted with a seriousness proportionate to the decision it may trigger.
That does not mean waiting forever before acting.
It means matching certainty to evidence.
New evidence should change the model. Good calibration is knowing how much change the evidence has earned.
That is very different from treating every score as a verdict.
