Wait, What? The Work Can Stay the Same While the Judgement Moves
A teacher reads an average essay after marking five weak essays. It can look surprisingly strong. The same essay read after five excellent essays can look disappointing.
The student’s work has not changed. The frame around the judgement has.
This matters because school assessment is not performed by a disembodied scoring machine. Teachers compare, remember, infer and contextualise. Those capacities make expert judgement useful. They can also create halo and frame-of-reference effects that pull a score away from the evidence in the piece of work itself.
Quick Answer
Owned Bolt job: calibrate teacher judgement when surrounding performances, prior information or a general impression can shift how the same student work is evaluated.
Teacher judgement is often reasonably accurate and educationally essential. But judgement is context-sensitive. Recent experimental evidence shows that identical writing can receive different evaluations depending on the achievement level of the surrounding reference group. Other experiments show halo effects, where an earlier strong or weak performance influences judgement of a later average performance. Clear criteria, anchor exemplars, structured scoring and deliberate separation of dimensions can reduce—but not eliminate—these effects.
Two Different Context Effects
Frame-of-reference effect
A performance is judged relative to the quality of other performances currently in view. An average script may appear stronger in a weak set and weaker in a strong set.
Halo effect
A judgement about one characteristic or earlier performance spills into another judgement. A student known as “strong” may receive the benefit of ambiguity; a student carrying a “weak” label may have an equally ambiguous response read more harshly.
These processes are related but not identical. Both create the same Bolt problem: is the mark following the current evidence, or is prior context pulling the mark?
Recent Evidence: Context Really Can Move the Assessment
A 2025 Learning and Instruction study directly tested frame-of-reference effects in teachers’ assessment of writing. In experimental and field settings, the same quality of student text could receive more positive evaluations when embedded in a lower-achieving reference context and more negative evaluations in a higher-achieving context. The effect appeared among both experienced and student teachers.
An experimental study of halo effects in grading found that an earlier weak, average or strong performance influenced the grade given to a later performance that was held constant at an average level. In other words, what the marker had already learned about the student changed the later judgement.
At the same time, a 2024 psychometric meta-analysis of teacher judgement accuracy found that teachers’ achievement judgements contain substantial real signal and that earlier reviews may actually have underestimated accuracy after accounting for measurement artefacts.
Those findings belong together. Teacher judgement is neither worthless nor perfectly objective. It is useful professional evidence that improves when its context effects are actively managed.
Why Expertise Does Not Automatically Remove the Problem
Experienced teachers know more about subject quality, developmental progression and common errors. That expertise can improve judgement. It also means teachers carry richer prior models of students and classes.
When time is short, those models can act as efficient heuristics. A teacher may think, “This student usually reasons carefully, so this phrase probably means…” Such interpretation can sometimes be correct. But if the task is to score the current performance against explicit criteria, the same prior model can become a source of bias.
Research on teacher diagnostic reasoning makes this trade-off visible: holistic impressions can be functional, but assuming consistency between achievement, motivation and other learner characteristics can also produce halo effects.
School, Teacher and Student: Three Different Consequences
School
A school should design marking and moderation so important decisions do not depend entirely on uncontrolled comparison effects. Shared criteria, exemplars, moderation, second marking where stakes justify it, and careful ordering of scripts can improve score stability.
Teacher or Coach
The teacher should separate diagnostic knowledge of the learner from evidence in the current performance. Prior knowledge is useful when deciding what to teach next. It should have less authority when the job is to score a particular script fairly against declared criteria.
Student
A student should know that one mark is evidence, not an oracle. If a judgement appears surprising, the productive response is not “the teacher is biased” by default. It is to ask which criterion was met, where the evidence appears in the work, and whether another calibrated reading agrees.
Competing Explanations for an Unexpected Mark
- The work genuinely met the criteria differently from the student’s or parent’s expectation.
- The marker’s frame of reference shifted after reading unusually strong or weak scripts.
- Prior knowledge of the student influenced interpretation of ambiguous evidence.
- The rubric itself was too vague to constrain judgement sufficiently.
- The performance sits near a boundary where legitimate professional disagreement is common.
- The marker changed standard over time.
- A second marker is using a different but defensible interpretation of the criteria.
Several of these have their own Bolt owners. This page owns only the context carried into the judgement.
The Bolt Context-Control Calibration Protocol
- Name the scoring construct. What dimensions of this performance are actually being judged?
- Use explicit criteria. Replace global impressions with observable features where the construct permits it.
- Anchor the scale. Use agreed exemplars at important performance levels.
- Score dimensions before forming the global impression. This can reduce the chance that one feature dominates all others.
- Control script order where stakes matter. Avoid letting one unusual cluster become the accidental standard for everything that follows.
- Blind irrelevant prior information when appropriate. The marker may not need previous grades, class labels or teacher comments to score the current script.
- Moderate samples, not only borderline cases. Shared standards are easier to maintain when teachers regularly compare reasoning.
- Investigate large discrepancies. Ask which criterion produced the disagreement rather than simply averaging marks.
- Re-score anchors over time. This distinguishes context effects from longer-term rater drift.
- Keep diagnosis separate from scoring. After the score is secured, prior knowledge can return to help interpret what teaching should happen next.
Worked Example: The Average Essay Between Two Extremes
Two teachers independently score the same average essay. Teacher A has just marked six very weak scripts. Teacher B has just marked six excellent scripts.
Teacher A gives the essay 18/25. Teacher B gives 15/25. In moderation, both can justify their initial impression, but when they return criterion by criterion to common anchor scripts, they converge on 16/25.
The useful conclusion is not that the first marks were fraudulent. The judgement process was context-sensitive. Anchors restored the intended reference frame.
That is what calibration is for: not to eliminate professional judgement, but to keep it answerable to shared evidence.
Objective Criteria Help—but They Do Not Make Judgement Mechanical
A 2024 experimental study found that objective assessment criteria reduced the influence of judgemental bias on grading. This supports the practical value of clearer criteria.
But criteria themselves still require interpretation. Complex writing, oral performance, scientific explanation and creative work cannot always be reduced to a checklist without losing the construct. The goal is not to remove professional judgement. The goal is to structure it so irrelevant context has less room to dominate.
The Dangerous Shortcut: “I Know This Student”
Knowing a learner well is one of a teacher’s great advantages. It helps detect unusual performance, choose feedback, recognise effort and understand developmental history.
But the same knowledge should not become permission to overwrite current evidence. A normally excellent student can submit poor work. A normally weak student can produce excellent work. Calibration requires the teacher’s model of the learner to remain correctable by the latest performance.
Common Misconceptions
- “Teacher judgement is subjective, so tests are always better.” False. Tests have their own sampling, scoring and validity limits.
- “Experienced teachers are immune to context effects.” Recent research found frame-of-reference effects among experienced as well as student teachers.
- “Anonymous marking solves bias.” It can remove some irrelevant information but cannot remove script-order, rubric or frame effects.
- “Moderation means forcing everyone to the same opinion.” Good moderation makes the evidence and criterion reasoning visible.
- “A halo effect proves prejudice.” Halo is a cognitive judgement process; its presence does not by itself identify motive.
How Do We Know?
The 2025 Learning and Instruction article Context counts: Unveiling the impact of achievement level on teachers’ text assessment found frame-of-reference effects in both experimental and field studies: identical-quality text could receive different assessments depending on surrounding achievement context.
The experimental study Halo effects in grading: an experimental approach found that an earlier weak, average or strong performance influenced grading of a later average performance by the same nominal student.
The 2024 study Objective assessment criteria reduce the influence of judgmental bias on grading provides direct evidence that more objective criteria can reduce some contextual judgement bias.
For the broader accuracy question, the 2024 PLOS ONE article Teachers’ judgment accuracy: A replication check by psychometric meta-analysis found that teacher judgements contain substantial achievement signal and that previous meta-analyses may have underestimated accuracy after accounting for artefacts.
The evidence boundary matters: these effects vary by task, rubric, marker, context and design. There is no universal number of marks to subtract for “halo.” The appropriate response is procedural calibration, not a correction formula.
For Parents: Ask for the Criterion, Not a Character Judgement
If a mark seems surprising, the strongest conversation is specific: Which criterion was missed? What evidence in the work supports the score? What would a stronger version look like? Is there an anchor or exemplar?
That keeps the discussion on performance instead of turning one mark into a story about whether the teacher likes the child or whether the child “is” a certain kind of student.
Bolt Direction Graph
Student work → surrounding context/prior model → teacher judgement → criteria and anchor check → isolate current evidence → moderation where needed → secure score → return prior learner knowledge for teaching interpretation → recalibrate.
Useful neighbours: When Two Good Teachers Give Different Marks, The Marker Can Drift Even When the Rubric Does Not, and Student Feedback About Teaching Is Evidence, Not a Verdict.
