Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

Bolt Measurement Note 19 — The Marker Can Drift Even When the Rubric Does Not

Wait, What? The Same Teacher Can Mark Differently in March and October Without Realising It

A marking rubric is printed. The descriptors have not changed. The teacher is experienced. Yet the practical standard applied to student work can still shift over time.

Raters are human measurement instruments. They become more familiar with the rubric, encounter stronger and weaker scripts, adjust expectations, tire, recalibrate, or gradually develop a slightly different interpretation of what “good enough” looks like. In formal assessment this is called rater drift. In ordinary school life, it can happen quietly.

Quick Answer

Owned Bolt job: calibrate score interpretation when the human marking standard itself may change across time.

If a teacher marks more severely or leniently later in the year, score change can partly reflect a change in the rater rather than a change in student performance. The correct response is not to distrust teachers. It is to recognise that stable criteria require active calibration, especially when marks are compared across time or used for high-stakes decisions.

Why a Rubric Does Not Automatically Hold the Standard Still

A rubric provides language, categories and anchors. It does not eliminate judgement. Two problems can remain:

  • Between-rater variation: two teachers apply the same rubric differently.
  • Within-rater drift: the same teacher changes how the rubric is applied over time.

Bolt Measurement Note 03 already owns the first problem: two good teachers can give different marks. This page owns the second. The measurement instrument itself can move.

What Can Make a Marker Drift?

  • Repeated exposure to very strong or very weak work can shift the internal reference point.
  • Rubric language can gradually be interpreted more strictly or loosely.
  • New examples may redefine what the marker considers typical.
  • Fatigue or workload can alter attention to criteria.
  • Feedback from moderation may recalibrate the rater.
  • Teachers may become more expert at noticing features they previously overlooked.
  • Institutional expectations may change even when the printed rubric remains unchanged.

Some of these changes are improvement. A rater who becomes better calibrated should change. The measurement issue is whether score comparisons across time still mean what users think they mean.

School, Teacher and Student: The Drift Problem Looks Different From Each Position

School

If scores are used to track progress, compare classes or make placement decisions, the school should ask whether marking standards are sufficiently stable across raters and time. Moderation, anchor scripts and periodic recalibration are not bureaucratic extras when score comparability matters.

Teacher or Coach

A teacher should periodically re-check marking against shared examples rather than assuming consistency because the rubric is familiar. The goal is not mechanical conformity. It is to make sure that a score change in the student is not partly a hidden score change in the marker.

Student

If a student’s writing score falls from 18/25 to 15/25, that matters. But before deciding that writing has deteriorated, compare the task, rubric, marker and moderation conditions. A changed mark can reflect a changed performance, a changed task, a changed standard—or several at once.

The Bolt Rater-Drift Calibration Protocol

  1. Preserve anchor work. Keep a small set of previously agreed scripts representing important score points.
  2. Re-score periodically. Ask whether the current judgement still matches the earlier calibrated standard.
  3. Moderate disagreements. Discuss why the criterion was applied differently rather than merely averaging marks.
  4. Record rubric changes. If interpretation legitimately changes, make the change visible instead of pretending the scale remained identical.
  5. Watch severity over time. Look for systematic movement toward harsher or more lenient scoring.
  6. Separate student trend from rater trend. When progress claims are important, inspect whether marking conditions were comparable.
  7. Increase calibration when stakes rise. High-stakes decisions justify stronger moderation than low-stakes daily feedback.

Worked Example: The Essay Standard Tightens

In Term 1, a teacher awards 17/25 to a set of essays. By Term 4, after months of reading stronger scripts and participating in moderation, the teacher has become more demanding about evidence and precision. A Term 1-style essay would now receive 14/25.

The teacher may genuinely have improved as a marker. But if the school charts student progress using raw marks across the year, it can mistakenly conclude that students have stalled. A better system re-scores anchor scripts or moderates across time so the scale itself remains interpretable.

This is the heart of Bolt calibration: before interpreting movement in the learner, check whether the measurement frame moved too.

What This Does Not Mean

  • Human marking is not useless. Many important performances require expert judgement.
  • Perfect agreement is not the goal. Two raters can agree and still share the same bias.
  • Drift is not misconduct. It is a measurement risk that can occur even among conscientious professionals.
  • Moderation does not mean every teacher must think identically. It means consequential score interpretations need a defensible shared standard.
  • A fixed rubric is not sufficient evidence of a fixed scale. Application matters.

How Do We Know?

Rater drift is a recognised problem in educational measurement. The 2024 Applied Measurement in Education article New Tests of Rater Drift in Trend Scoring develops statistical methods specifically for detecting changes in rater scoring across occasions. Its focus is formal trend scoring, but the underlying measurement lesson is widely relevant: a human scorer’s applied standard can change over time and should be monitored when comparability matters.

The 2024 study Signal, error, or bias? exploring the uses of scores from observation systems provides a broader reminder that rater error and systematic bias can materially affect educational observation scores, while Bolt Measurement Note 03 addresses between-rater disagreement in ordinary marking.

The evidence boundary matters: classroom teachers usually do not need formal psychometric drift statistics. For school use, anchor examples, moderation and periodic re-checking can provide a practical approximation. The more consequential the decision, the stronger the calibration procedure should be.

For Parents: A Falling Mark Is a Question, Not Yet an Explanation

If a child’s score falls even though the work looks stronger, ask whether the task or marking standard changed. Do not assume unfairness, and do not assume decline. Ask for the rubric, an example, and what changed in the expected standard.

Sometimes the answer will be: the learner needs to improve. Sometimes it will be: the scale became stricter. Good calibration can hold both possibilities open until the evidence separates them.

Bolt Direction Graph

Scored performance → preserve rubric and anchors → re-score across time → detect severity/leniency shift → moderate → separate learner change from marker change → recalibrate progress claim.

Useful neighbours: Bolt Measurement Note 03 — When Two Good Teachers Give Different Marks, Bolt Measurement Note 02 — Before You Call It Improvement, Check Whether the Scores Are Comparable, and Bolt 29 — When the Evidence Disagrees.