Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

Bolt Measurement Note 03 — When Two Good Teachers Give Different Marks

Three students studying together in an eduKate small-group classroom.

Bolt Measurement Notes · Supplementary to the Bolt 01–40 Core · Note 03

Wait, What? Two competent teachers can look at the same work and disagree

A student writes one essay.

Teacher A gives 72.

Teacher B gives 78.

Which teacher is wrong?

Sometimes one is.

Sometimes the marking criteria were interpreted differently.

Sometimes the task permits legitimate professional judgement and the disagreement sits inside a reasonable range.

Sometimes the real problem is not either teacher. It is that the assessment system has not made its standard sufficiently shared, visible or stable.

Quick Answer

Human judgement is evidence, not magic. Teacher expertise can add information that a simple machine score cannot see, but human marking can also vary. Good assessment systems therefore do not solve the problem by pretending judgement is perfectly objective. They train assessors, clarify criteria, use exemplars, standardise, moderate, quality-check and review difficult cases.

For Bolt, disagreement is not automatically proof that assessment is useless. It is a calibration signal:

How much of this difference belongs to the learner’s performance, and how much belongs to the measurement process?

Owned Calibration Job

This article owns one narrow Bolt problem:

How should schools, teachers, parents and students interpret performance when human assessors may apply judgement differently?

This is not the same as Bolt 15 — When the Coach Is Wrong, which concerns coach judgement more broadly. Here the object is the measurement process itself: marking consistency, shared standards, moderation and what disagreement does to confidence in a score.

It does not own the learner’s next study operation or the concrete feedback task. Those remain with MindOS and the Student/Studying Interface.

A mark contains at least two things

When a human assessor marks a response, the final result reflects at least:

  • the evidence actually present in the student’s work; and
  • the assessor’s application of criteria to that evidence.

In highly constrained items, the second component may be small. Two plus two either equals four or it does not.

But many important educational performances are not binary.

  • How well does this essay develop an argument?
  • How convincing is the scientific explanation?
  • How effectively does the student select and use evidence?
  • How sophisticated is the interpretation of a text?
  • How complete is the reasoning in a multi-step solution?
  • Does the practical performance meet a standard?

Those questions require judgement.

The answer is not to remove judgement from education. The answer is to calibrate it.

Teacher judgement is useful — and imperfect

Research on teacher judgement shows an important middle position.

Teachers are not random observers. A recent psychometric meta-analysis found a substantial association between teacher judgements and student achievement, while also showing why measurement error, sampling and publication bias matter when estimating judgement accuracy.

Earlier review work likewise found that teacher judgements can be reasonably accurate on average while accuracy varies by task, information available, learner characteristics and the way judgement is elicited.

That means “trust teachers” and “teachers are subjective” are both too crude.

Professional judgement becomes stronger when the evidence, criteria and calibration process around it become stronger.

Why professional systems standardise markers

Ofqual’s 2025 delivery report describes a deliberate quality-control process for external examinations in England. Examiners are trained on the mark scheme before marking. Their work is quality checked during marking. Examiners who do not apply the agreed standard can be stopped, retrained or have previous marking reviewed.

For internally marked assessments, schools and colleges standardise marking among teachers, while awarding organisations moderate samples to check that the required standard is being applied accurately and consistently.

These systems exist because human judgement can be valuable and variable.

Reliability does not require pretending humans never disagree. It requires enough consistency for the intended interpretation and decision.

Five different kinds of marking disagreement

1. A simple marking error

A correct answer is missed, a total is added wrongly, or a rubric rule is applied incorrectly. This is the clearest case.

2. Different interpretations of the criterion

One teacher interprets “well-developed explanation” more strictly than another. Standardisation and exemplars can help align the threshold.

3. Different weighting inside a holistic judgement

One marker gives more weight to conceptual sophistication; another reacts more strongly to organisation or technical accuracy. If the rubric does not define the balance clearly, disagreement may be structural rather than personal.

4. Borderline evidence

The work may genuinely sit near a threshold. Small judgement differences can then move the final mark or grade even when both assessors are broadly aligned.

5. The construct itself is difficult to reduce to one number

Complex performances such as writing, design, oral communication or practical work often contain multiple dimensions. A single score compresses them. Some disagreement can therefore reveal that the underlying performance needs higher-resolution description.

Do not treat disagreement as a student defect

If one teacher gives 72 and another gives 78, the student did not suddenly become six points different.

The work stayed the same.

That distinction matters emotionally and educationally.

The learner should not be told:

“You are a 72 student.”

Nor should the higher mark automatically be selected because it feels nicer.

The calibration question is whether the disagreement changes the conclusion we are entitled to draw.

The moderation test: can assessors explain the evidence against a shared standard?

Moderation is strongest when it is not simply “average the two marks.”

A useful professional conversation asks:

  • Which exact evidence in the work supports this judgement?
  • Which criterion is being applied?
  • What does an agreed exemplar at the boundary look like?
  • Where do the two assessors actually disagree: evidence, criterion, threshold or weighting?
  • Would the same reasoning be applied to another student’s work?
  • Does the final judgement remain consistent with the assessment’s intended purpose?

This moves the discussion away from authority—“I have taught longer, therefore I am right”—and toward inspectable evidence.

For schools: design calibration before the dispute

Moderation is easier when schools build shared standards before marks become consequential.

  • Use clear criteria connected to the intended construct.
  • Discuss annotated exemplars before large marking exercises.
  • Calibrate on a small common sample.
  • Re-check marker consistency during long marking periods.
  • Escalate borderline or high-stakes cases appropriately.
  • Preserve evidence of why a difficult judgement was made.

For formal qualifications, follow the applicable awarding and regulatory procedures. Classroom calibration practices are not substitutes for official moderation systems.

For teachers: separate feedback from scoring certainty

A teacher can be highly confident that an essay needs a clearer argument while being less certain whether it deserves 71 or 74.

Those are different levels of inference.

The educationally useful feedback may be stable even when the exact mark has some uncertainty.

This is important because students can spend enormous energy arguing over one mark while overlooking the stronger signal in the work.

Bolt therefore asks two questions:

  • How certain are we about the score or level?
  • How certain are we about the performance feature that should be addressed next?

They may have different answers.

For students: disagreement can teach you measurement literacy

If two teachers disagree, do not immediately search for the teacher who gives the higher score.

Ask:

  • What evidence did each teacher use?
  • Which criterion created the difference?
  • Is there an exemplar showing the expected standard?
  • What part of the feedback appears in both judgements?
  • What later performance would settle the uncertainty?

The overlapping feedback is often more useful than the fight over the exact number.

What review statistics do — and do not — tell us

Ofqual reported that in summer 2025 a minority of all GCSE, AS and A level grades were challenged, and an even smaller share of all grades changed following review. Among the subset that was challenged, however, some reviews did lead to changed grades.

Those figures should not be turned into a universal “error rate.” Reviews are not a random sample of all marking, subjects differ, and a grade change is not identical to a marking error.

The useful lesson is narrower: mature assessment systems anticipate that review and correction mechanisms are sometimes necessary.

Common Misconceptions

“If markers disagree, the assessment is invalid.”
Not automatically. Reliability is one part of validity, and acceptable consistency depends on purpose and stakes. But material disagreement requires investigation.

“Rubrics remove judgement.”
No. Good rubrics constrain and structure judgement; they do not necessarily eliminate it.

“The experienced teacher must be right.”
Experience is relevant, but calibration should return to criteria and evidence rather than status alone.

“Just average the two marks.”
An average can conceal a systematic standards problem. First understand why the marks differ.

How Do We Know?

Educational measurement guidance treats inconsistency in human marking as a reliability issue. Ofqual’s General Conditions explicitly identify human assessor inconsistency as one factor affecting reliability and require arrangements that support accurate, consistent application of assessment criteria.

The current Ofqual guide for schools describes standardisation within centres and moderation by awarding organisations for relevant internally marked assessments. Its 2025 delivery report details examiner training and ongoing quality checks for external marking.

Research on teacher judgement supports the value of informed professional observation while also demonstrating that judgement accuracy is not perfect and should be studied with attention to measurement error, sampling and context.

Evidence Boundary

Reliability is task- and purpose-dependent. Highly structured objective items, essays, performances and practical assessments create different marking demands. Evidence from one context should not be mechanically transferred to another.

  • Do not infer that teacher judgement is generally unreliable because one marking dispute occurred.
  • Do not assume a formal moderation procedure guarantees perfect agreement.
  • Do not turn a mark difference into a claim about teacher bias without evidence.
  • Do not diagnose learner problems from assessor disagreement.

The Bolt → Interface → MindOS → Bolt Handoff

Bolt calibrates: “The exact score is uncertain because assessors differ, but both identify weak evidence selection and strong conceptual understanding.”

The Student/Studying Interface makes that conclusion operable: the next assignment exposes the relevant criterion, exemplar, required evidence and first action. The Feedback Handoff turns the calibrated signal into something the learner can act on.

MindOS runs the learner operation: comparison, explanation, strategy selection or another learning operation is chosen only if the evidence supports it.

Return to Bolt: a new independently marked performance tests whether the targeted feature improved and whether assessor agreement is stronger.

Parent and Tutor Guide

If a child receives unexpectedly different judgements, avoid telling them either “marks are meaningless” or “the teacher is always right.”

Ask for the criterion, the evidence in the work and the standard being applied.

“What part of the judgement is stable even if the exact mark is not?”

That question protects respect for teachers while preserving the learner’s right to understand how evidence became a judgement.

Bolt Direction Graph

STUDENT WORK
→ criteria + evidence
→ assessor judgement A / B
→ locate disagreement
→ marking error? criterion? threshold? weighting? borderline evidence?
→ standardise / moderate where appropriate
→ bound score certainty
→ preserve stable performance findings
→ Interface translates finding into next task
→ MindOS runs justified operation
→ new performance
→ Bolt recalibrates

Authoritative Sources and Further Reading


Durable Bolt rule: A human mark is strongest when the judgement can be traced back to shared criteria, visible evidence and a calibration process that can detect disagreement rather than hide it.