Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

Bolt 03 — What Exactly Did This Test Measure?

Three students studying together in an eduKate small-group classroom.

Bolt Series · Human Performance Calibration · Article 03

A test can be marked perfectly and still answer a smaller question than you think

Here is a strange possibility.

A teacher can write a good examination.

A student can sit it fairly.

Every answer can be marked correctly.

Every mark can be added correctly.

And the final score can still be an incomplete description of the student’s capability.

Nothing has necessarily gone wrong.

The problem might simply be that we are asking the score to tell us something it was never designed to tell us.

This is where educational measurement becomes much more interesting.

Correct marking and correct interpretation are different problems

Suppose a student obtains 72%.

The arithmetic question is easy:

Did we add the marks correctly?

But there is another question:

What does 72% allow us to conclude?

That question is harder.

Modern testing theory treats validity as concerning whether evidence supports particular interpretations and uses of scores.

In other words, we do not merely ask, “Is this a valid test?” We also ask, “Valid for what conclusion and for what purpose?”

The National Council on Measurement in Education makes this explicit: understanding the intended purposes and uses of educational tests is part of the foundation for evaluating validity.

That single idea can prevent many educational mistakes.

Consider a 100-metre race

A fully automatic clock is extraordinarily good at answering:

How long did this athlete take to complete this 100-metre race?

It can also help decide who crossed fastest.

But imagine using the same result to answer:

  • Who would make the best marathon runner?
  • Who has the strongest legs?
  • Who has the best tactical judgement?
  • Who has the best cardiovascular health?
  • Who will be the best coach?

Now the stopwatch has not become inaccurate.

Our inference has become unreasonable.

Usain Bolt’s 9.58 seconds is highly meaningful for elite 100-metre sprint performance. It does not measure every interesting property of Bolt.

Accurate instruments can have narrow jobs.

School tests can too.

Every test samples

Imagine a student has learned a large Mathematics syllabus.

An examination cannot ask every possible question.

It samples.

It might contain three algebra questions, two geometry questions, a statistics problem, a difficult application question and several routine procedural items.

Change the sample and the score may change.

That does not make assessment useless. Sampling is unavoidable.

ETS guidance on reliability notes that no test can be perfectly reliable because a test is a sample from a larger population of possible questions, possible times and possible conditions of administration. Observed scores therefore contain measurement error.

The important lesson is simply:

The score describes performance on a sample.

We should think before extending that conclusion much further.

Weighting matters too

Suppose two examinations cover exactly the same topics.

Paper A gives 40% of its marks to algebra.

Paper B gives algebra 15%.

A student who is exceptionally strong in algebra and weaker in geometry may produce different overall scores.

Which score is correct?

Potentially both.

They answer slightly different weighted questions about performance.

This is why assessment design matters.

The final number does not descend from the sky. Someone chose what to test, how often to test it, how difficult the questions should be, how marks should be distributed, what counts as evidence of understanding, how much time is allowed and how responses are scored.

Good assessment design makes those choices carefully.

But they remain choices.

State can enter the measurement too

Now add the human being.

A student’s observed performance may interact with sleep, attention, stress, language, time pressure, motivation, health, familiarity with the format and normal day-to-day variability.

We should not use these possibilities as automatic excuses whenever a score disappoints us.

That would destroy calibration in the opposite direction.

Instead, we ask a scientific question:

Is this performance consistent with the wider evidence?

If a student has repeatedly demonstrated the same weakness, the examination may be confirming something important.

If one result sits dramatically outside months of other evidence, perhaps we should investigate before making a large conclusion.

The answer comes from additional evidence.

Not denial.

Fair measurement requires intellectual humility

There is a subtle but important difference between saying:

“This test is unfair because I didn’t like my result.”

and:

“What inference does this assessment genuinely support?”

The first can become an excuse.

The second is measurement literacy.

Good measurement is not weak because it recognises uncertainty.

It is stronger.

Why this matters in education

We frequently ask scores to do too many jobs.

We want one examination to tell us whether the child understood, whether they studied, whether they are intelligent, whether the teacher taught well, whether they will succeed next year, whether they belong in a particular class, and sometimes even whether they are “good” or “bad” at the subject.

That is an enormous amount of inferential weight to place on one number.

A better approach is to ask:

  • What did this assessment genuinely measure well?
  • What did it measure less well?
  • What additional evidence would I want before making an important decision?

That is not anti-examination.

It is pro-measurement.

A good test deserves a good interpretation

Tests are most useful when we respect their boundaries.

A ruler measures length extremely well. We do not criticise it for failing to measure temperature.

A stopwatch measures time extremely well. We do not ask it to measure courage.

And an examination can provide highly useful evidence about particular knowledge and performance without becoming a complete description of the learner.

So perhaps the question after every important score should not simply be:

“What did you get?”

It should also be:

“What exactly did this test measure?”

That question changes everything that comes next.

Evidence and further reading