Bolt Series · Human Performance Calibration · Article 27
How many performances do we need before we believe the pattern?
One test can be noisy.
So we say:
“We need more evidence.”
Fair enough.
But how much more?
Two tests?
Five?
Twenty?
There is no universal number.
That may be frustrating.
It is also scientifically important.
The amount of evidence needed depends on what we are trying to infer, how noisy the measurement is, how much the tasks vary, and how consequential the decision will be.
Reliability is partly a question of repeatability
In assessment, reliability concerns the consistency of scores.
If the same underlying capability is measured again under reasonably comparable conditions, how much should we expect the result to change?
If the measurement is highly unstable, one reading deserves limited confidence.
If repeated readings converge, confidence can increase.
Research on performance assessment uses generalizability theory to separate sources of variation such as tasks, raters, occasions and the person being assessed.
That research shows why a single observation often has limited power when performance varies across cases or observers.
It also shows why adding observations can increase reliability.
More observations help when they add genuinely new information rather than merely repeating the same narrow sample.
Five copies of the same task are not the same as five independent views
Imagine a student completes five nearly identical algebra exercises.
They score perfectly on all five.
We have repeated evidence of procedural fluency on that form of the task.
But we still do not know whether the student can:
- retrieve the method after a delay;
- recognise when the method applies;
- transfer it to an unfamiliar representation;
- explain why the method works;
- perform it under examination time pressure.
Five repetitions can increase confidence about one narrow claim.
They do not automatically increase confidence about every larger claim.
Evidence becomes stronger when the pattern survives variation
Suppose the same capability appears across:
- different question forms;
- different days;
- different levels of support;
- different contexts;
- different raters where judgement is involved;
- different levels of time pressure.
Now we have something more powerful.
The pattern is not only repeating.
It is surviving changes in the environment.
That matters because capability is often something we want to generalise beyond the exact task used to measure it.
Modern educational measurement explicitly treats generalisation from the observed sample to a wider universe of possible questions or performances as part of the validity argument.
Mechanism can sometimes matter more than count
Imagine a student makes the same sign error in three different algebra problems.
The contexts differ.
The surface of the questions differs.
But each failure occurs at the exact same operation.
That repeating mechanism may be more diagnostically convincing than ten unrelated wrong answers.
This gives teachers another principle:
Do not count errors only. Look for the thing that repeats inside the errors.
The stakes should determine how much certainty we demand
Not every educational decision needs the same evidence threshold.
Choosing tomorrow’s practice set is a low-cost decision.
We can act on a plausible diagnosis and revise quickly if the next performance disagrees.
A decision that changes a student’s long-term pathway, access to opportunity or educational placement deserves much stronger evidence.
The higher the consequence, the more we should care about reliability, validity, sampling and alternative explanations.
This is one reason good measurement is not merely about collecting data.
It is about matching the strength of the conclusion to the strength of the evidence.
One excellent performance can still be important
Repeated evidence is valuable, but we should not become blind to exceptional observations.
Suppose a student who has always been given routine work suddenly solves a genuinely difficult unfamiliar problem independently.
That single performance does not prove that every advanced skill is secure.
But it may decisively weaken the belief that the student is incapable of higher-level reasoning.
Evidence does not have equal informational value merely because each item counts as “one observation.”
A well-chosen discriminating task can reveal something that twenty routine tasks never tested.
So when should we trust a pattern?
Not when an arbitrary number has been reached.
Trust should rise when several conditions begin to converge.
- Repeatability: similar outcomes recur.
- Variation: the outcome survives meaningful changes in task or context.
- Mechanism: we can explain what is producing the pattern.
- Measurement quality: the instrument is appropriate and sufficiently reliable for the claim.
- Triangulation: learner report, external observation and performance receipts do not wildly contradict one another—or the contradiction itself has been investigated.
- Prediction: the model successfully predicts what happens next.
The last condition is especially powerful.
If we believe a learner’s weakness is transfer, we should be able to predict that routine work will succeed while unfamiliar applications continue to fail.
If that prediction repeatedly holds, confidence in the model rises.
If it does not, the model needs revision.
The learner can learn this evidence rule too
A student should not decide “I know this” because one answer worked.
Nor should they decide “I can never do this” because one attempt failed.
They can ask:
- Can I do it again?
- Can I do it later?
- Can I do it without help?
- Can I do it when the question changes?
- Can I explain why?
- Can I predict which version will still be difficult?
Now mastery is no longer a feeling.
It is an evidence claim.
Why this matters for education
Good calibration requires enough evidence to resist two temptations.
Overreaction:
“One result proves everything.”
And endless hesitation:
“We can never know enough to decide.”
Education needs a more useful middle.
Collect enough independent, relevant and repeatable evidence for the consequence of the decision—and keep the conclusion open to correction when better evidence arrives.
There is no magic number.
There is a better discipline.
