Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

Bolt Measurement Note 02 — Before You Call It Improvement, Check Whether the Scores Are Comparable

Three students studying together in an eduKate small-group classroom.

Bolt Measurement Notes · Supplementary to the Bolt 01–40 Core · Note 02

Wait, What? 78% can be worse evidence of improvement than 72%

A student scores 72% on one test and 78% on the next.

Improvement?

Possibly.

But before we congratulate, worry, reward, reteach or change the learner model, there is a more basic question:

Were those two scores measuring comparable performances?

If the second paper was easier, sampled different content, allowed more help, used a different marking scheme or placed less demand on the learner, a higher raw percentage does not automatically mean capability increased.

Equally, a lower score on a harder, broader or more independent assessment does not automatically mean the learner went backwards.

Quick Answer

Change in score is not the same thing as change in capability unless the score meanings are sufficiently comparable for the conclusion you want to make.

Large testing programmes use equating, standard setting, moderation and other technical procedures because different forms cannot simply be assumed to have identical difficulty and meaning. In ordinary classrooms, teachers rarely need that level of psychometric machinery. But they do need the underlying discipline: compare like with like, record important condition changes, and treat unmatched scores as evidence with uncertainty rather than a perfect growth ruler.

Owned Calibration Job

This article owns one narrow Bolt job:

When two performances produce different scores, are the measurements comparable enough to justify saying the learner improved, declined or stayed stable?

It does not own test design as a whole. It does not own what the student should practise next. It does not own examination technique. It does not turn classroom quizzes into formal psychometric instruments.

Its job is narrower: protect the inference of change from measurements whose meanings changed underneath us.

The hidden assumption inside every progress chart

A graph can make progress look wonderfully precise.

62 → 68 → 73 → 79.

The line rises. The story feels obvious.

But the graph contains a hidden assumption: that a point on the vertical axis means roughly the same thing each time.

That assumption may be reasonable when the assessments are deliberately built and interpreted to support comparison. It may be much weaker when one number came from a short open-book worksheet, another from a difficult timed test, another from group work and another from a heavily rehearsed practice paper.

The numbers are real. The line may still be misleading.

Professional assessment systems take comparability seriously for a reason

ETS describes score equating as essential when a testing programme produces new editions of a test but expects reported scores from those editions to have the same meaning over time. Different forms inevitably differ to some extent, so equating is used to support fair comparison.

Ofqual makes a similar distinction in regulated qualifications. Its guidance treats comparability as necessary when outcomes from different forms are used to compare learners or standards over time. In England, grade boundaries can change from year to year partly because examination papers differ in difficulty; the intention is to maintain a comparable standard of work rather than pretend that the same raw mark must always mean the same thing.

That does not mean classroom teachers should start equating every spelling test or homework quiz.

It means the measurement principle is real:

A number can only be compared as confidently as its measurement conditions allow.

Six ways the meaning of a score can change

1. The questions changed in difficulty

A paper with more routine items is not equivalent to a paper requiring more unfamiliar transfer, even if both contain 50 marks.

2. The content sample changed

A student can improve in algebra while the next assessment places more weight on geometry. The total score may fall while an important capability rises.

3. The support conditions changed

Open notes, teacher prompts, calculators, AI tools, peer discussion or visible worked examples can change the performance being measured. Bolt Measurement Note 01 therefore separates supported performance from independent performance.

4. The time conditions changed

A learner may demonstrate the same knowledge differently under generous time and under a strict time limit. Whether that difference should count depends on what the assessment intends to measure.

5. The marking changed

Different rubrics, stricter interpretation, partial-credit rules or human marker judgement can move scores even when the underlying work is similar.

6. The student changed strategy for the test rather than capability itself

Familiarity with question format, memorised templates or repeated exposure to almost identical items can improve a score without guaranteeing broader transfer.

A higher score can still be good news

Measurement caution should not become paralysis.

If a learner repeatedly scores higher on assessments with similar purpose, coverage, difficulty, support rules and marking standards—and the improvement survives new questions and later checks—that is increasingly persuasive evidence of progress.

We are not trying to make everyday education perfectly controlled.

We are trying to stop obvious condition changes from disappearing behind a neat percentage.

The classroom comparability check

Before saying “up six marks means improvement,” check five things.

  1. Construct: Were both assessments trying to measure substantially the same knowledge or performance?
  2. Demand: Was the level of difficulty and transfer demand reasonably similar?
  3. Conditions: Were time, tools, prompts and independence requirements comparable?
  4. Scoring: Were marking rules and standards stable enough for comparison?
  5. Replication: Does the apparent improvement appear again on another meaningful performance?

For low-stakes classroom decisions, a rough answer may be enough. For high-stakes placement, grading or major claims about progress, the evidence burden should rise.

Competing explanations for 72% → 78%

  • The learner genuinely improved the target capability.
  • The second paper was easier.
  • The second paper sampled the learner’s stronger topics.
  • The learner received more support.
  • The learner had practised very similar items.
  • The marking was more generous.
  • The learner was in a better temporary state.
  • The difference is ordinary performance or measurement variation.
  • Several of these occurred together.

Bolt does not choose the most comforting explanation or the harshest one. It asks what evidence can discriminate among them.

The discrimination test: build one common anchor

In a classroom, one practical way to improve comparability is to preserve a small common anchor across time.

That might be:

  • a recurring type of unseen problem at a stable difficulty;
  • a short set of common retrieval questions;
  • a rubric criterion scored the same way across assignments;
  • a repeated independent writing task with comparable demands;
  • a stable oral explanation prompt;
  • one transfer item that deliberately changes surface features while preserving the target concept.

This is not formal equating. It is a classroom calibration aid.

If the learner improves on both the changing curriculum test and the stable anchor, the progress story becomes stronger. If the overall score rises while the anchor does not, investigate before concluding that the underlying capability changed.

For schools: never make a longitudinal graph more precise than the assessments

School dashboards can create an illusion of exact continuity. A score from September, January and May may be placed on one smooth line even when the assessments differ substantially.

Before interpreting such a trend, schools should know:

  • whether the assessments were designed for longitudinal comparison;
  • whether forms were linked, equated, moderated or otherwise standardised;
  • whether curriculum coverage changed;
  • whether support and administration rules changed;
  • whether the score scale itself has a stable interpretation over time.

If those conditions are weak, the dashboard can still be useful—but the language should become more modest: “recent performance rose” rather than “capability increased by exactly six percentage points.”

For teachers: compare the work, not only the totals

Two total scores can hide important movement.

A learner may stay at 70% while:

  • moving from routine questions to unfamiliar transfer;
  • needing fewer prompts;
  • making stronger method selections;
  • writing better explanations;
  • becoming faster without losing accuracy;
  • retaining knowledge over a longer delay.

The unchanged total can coexist with real improvement because the performance demand changed.

The reverse can happen too: a higher score can conceal a narrower or more supported performance.

For students: do not let one harder test erase genuine progress

If your mark falls, ask what changed before turning the result into an identity statement.

  • Was the paper harder?
  • Did it include more unfamiliar transfer?
  • Was support removed?
  • Were more topics sampled?
  • Did the same old weakness actually reappear?
  • What stayed strong despite the harder conditions?

None of these questions are excuses. They are ways of interpreting the evidence accurately enough to choose the next action.

Common Misconceptions

“Higher percentage always means improvement.”
No. The score must have a sufficiently comparable meaning for that inference.

“If tests differ, scores are useless.”
No. They remain evidence of those performances. We simply bound cross-test claims more carefully.

“Same percentage means same standard.”
Not necessarily. Different papers can differ in demand, coverage and scoring.

“Classroom teachers need formal equating.”
Usually not. The practical lesson is to design sensible anchors, document material condition changes and avoid false precision.

How Do We Know?

Comparability is a core problem in educational measurement. ETS notes that score equating is required when multiple test forms are intended to produce scores with the same meaning over time. Its research also warns that imperfect samples, differing test properties and accumulated equating error can create score-scale inconsistency.

Ofqual’s regulatory guidance defines comparability as generating assessment outcomes that support fair comparison across forms and over time. Its current guidance for schools explains that grade boundaries can vary from year to year to account for differences in paper difficulty while maintaining a comparable standard of work.

These are formal assessment systems, so their technical procedures should not be copied casually into classroom practice. Their relevance here is conceptual: progress claims require some stability in what the numbers mean.

Evidence Boundary

No two classroom performances are perfectly identical. Perfect comparability is neither possible nor necessary for every educational decision.

The stronger the claim and the higher the stakes, the stronger the comparability evidence should be.

  • A small low-stakes teaching adjustment may tolerate rough comparison.
  • A claim that a learner has mastered a domain deserves broader evidence.
  • A placement, certification or high-stakes judgement demands formal assessment quality appropriate to the decision.

This article does not provide a psychometric equating procedure. It provides a calibration question for schools, teachers, parents and learners.

The Bolt → Interface → MindOS → Bolt Handoff

Bolt calibrates: “The second score is higher, but the paper was easier and support conditions changed. We do not yet have strong evidence that independent capability increased.”

The Student/Studying Interface makes that finding operable: the next task states a comparable target, matching support conditions, clear success criteria and a return point. The Progress Evidence Interface helps keep visible activity separate from justified progress claims.

MindOS runs the justified learner operation: only after the evidence locates the likely weak operation should the learner retrieve, compare, explain, practise, transfer or use another learning process.

Return to Bolt: a later performance under meaningfully comparable conditions tests whether the improvement reproduces.

Parent and Tutor Guide

When marks change, replace the first emotional question—“Why did you go up?” or “Why did you drop?”—with a measurement question:

“What was the same, and what changed between these two performances?”

Then inspect the actual work. If the better performance survives similar or harder conditions, celebrate the evidence. If the conditions were very different, design one fairer return check rather than arguing from the percentage alone.

Bolt Direction Graph

SCORE A → SCORE B
→ same construct?
→ comparable demand?
→ comparable support / time / tools?
→ comparable scoring?
→ stable anchor or repeated evidence?
→ bound the improvement claim
→ Interface creates a fair next performance
→ MindOS addresses the justified weak operation
→ new comparable performance
→ Bolt recalibrates

Authoritative Sources and Further Reading


Durable Bolt rule: Before treating a changing score as a changing learner, check whether the score meaning stayed stable enough to support the comparison.