Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

Bolt Measurement Note 23 — Human-Like Agreement Does Not Make an AI Marker Valid

Wait, What? An AI Can Agree With Human Markers Most of the Time and Still Be Unsafe for the Decision You Want to Make

Suppose an automated scoring system matches human marks on 90% of student responses. That sounds excellent. It may be excellent. But “agrees with humans” is only one piece of validation evidence.

The remaining questions are harder: Which responses produce disagreement? Are errors symmetric? Does the system behave differently across language backgrounds or response styles? Does it reward superficial features? Does performance hold when prompts change? Does the model still behave the same after an update?

Bolt treats automated marking as a measurement system, not a magic replacement for judgement.

Quick Answer

Owned Bolt job: calibrate when automated or AI-assisted scoring is sufficiently valid, reliable and fair for an educational decision.

High agreement with human scoring is useful evidence, but it does not prove full validity. A scoring system must also be tested for the construct it is supposed to score, error patterns, fairness, robustness across tasks and groups, operational stability, and the consequences of mistakes. Low-stakes formative use and high-stakes certification require different evidential burdens.

Agreement Is Not the Same as Validity

If two markers agree, several things are possible:

  • both are measuring the intended construct well;
  • both are using the same imperfect rubric;
  • the AI has learned patterns that correlate with human scoring without representing the intended construct;
  • agreement is strong overall but weak on important edge cases;
  • both systems share a systematic bias;
  • the agreement holds only for the dataset on which the system was validated.

This is why educational measurement does not reduce validity to one coefficient. A score has to support the interpretation and use being made from it.

The Better the Automation, the More Carefully We Must Ask What It Is Automating

Automated scoring can provide real benefits. It can reduce turnaround time, support large-scale assessment, provide consistency, and act as a quality-control layer. Recent work in international large-scale assessment shows that automated approaches can achieve impressive scoring performance under carefully designed conditions.

But the question is never simply “Can the model score?” The question is:

  • What response types has it been validated on?
  • What scoring construct is represented?
  • How close is its agreement to human-human agreement?
  • Where are the disagreements concentrated?
  • What happens to unusual but valid responses?
  • Does performance differ across demographic or language groups?
  • Can students game surface features that the model overweights?
  • What happens after the model, prompt, rubric or deployment pipeline changes?

School, Teacher and Student: Three Different Risks

School

A school deciding whether to use automated scoring should begin with stakes. A system used to provide a provisional practice score can tolerate a different error profile from a system used to determine progression, placement or certification. The higher the consequence, the stronger the need for human oversight, documented validation, appeals, subgroup analysis and operational monitoring.

Teacher or Coach

Teachers should understand what the automated score is good at detecting and where it is weaker. If the system is strong on surface language accuracy but weaker on originality, reasoning or discipline-specific nuance, the teacher should not treat the automated score as a complete representation of the response.

Student

Students need to know whether a machine score is provisional, formative or consequential, and whether there is a route for review. A student should not be forced to infer that an opaque automated judgement is infallible simply because it arrived instantly.

The Edge-Case Problem

Average agreement can hide concentrated failure. Imagine an AI scorer that agrees with human markers on almost every conventional response but systematically undervalues:

  • unusual but correct mathematical explanations;
  • responses using a less common language structure;
  • creative writing that deliberately breaks standard patterns;
  • partially correct reasoning expressed in an unexpected order;
  • very short but conceptually complete answers.

The total agreement statistic may remain impressive. The educational consequence for those students may still be serious. Bolt therefore asks not only “How often is the model right?” but “For whom, on what, and when is it wrong?”

Fairness Must Be Tested, Not Assumed From Accuracy

A high-performing model can still have subgroup-specific error patterns. A 2025 study of automatic short-answer scoring using 38,722 PISA reading-comprehension responses examined fairness across gender and home-language groups. The most accurate model showed no discernible gender disparity but did show a small significant language-background bias.

Another 2025 study examining GPT-4 short-answer grading found encouraging consistency across the groups studied, but the authors still called for further validation. These findings point in the same direction: fairness should be investigated empirically for the actual scoring system, task and population. It should not be inferred from the brand of model or from overall accuracy.

The Bolt Automated-Scoring Calibration Protocol

  1. Name the decision. Practice feedback, classroom marking, screening, placement and certification have different stakes.
  2. Name the construct. What exactly should the score represent?
  3. Establish human reference quality. If humans disagree badly, training the machine on their scores does not magically create truth.
  4. Measure more than overall agreement. Inspect false positives, false negatives, score severity, item-level errors and edge cases.
  5. Test subgroup fairness. Examine whether scoring error differs meaningfully across relevant groups.
  6. Challenge unusual valid responses. Search deliberately for examples the system may mishandle.
  7. Validate on unseen tasks and responses. Performance on development data is not enough.
  8. Monitor after deployment. Model, prompt, rubric and software changes can alter the scoring system.
  9. Keep a human escalation route. Especially when stakes are high or confidence is low.
  10. Recalibrate the use. A tool may be acceptable for triage or second-reader work while remaining unsuitable as the sole high-stakes judge.

Worked Example: The AI Essay Marker

A school pilots an AI marker for 500 essays. Agreement with teachers is high. The school is tempted to let the model produce final marks.

Before doing so, the assessment team examines disagreements. It discovers that the model is highly consistent on grammar and organisation but sometimes underrates essays that make sophisticated arguments using unconventional structures. Those essays are rare, so the overall agreement statistic barely changes.

The responsible conclusion is not “AI marking failed.” It is narrower: the model may be useful as a first pass or second reader, but the current evidence is insufficient for sole high-stakes scoring of this construct. Human review remains necessary for flagged or high-consequence cases.

What This Does Not Mean

  • Automated scoring is not inherently invalid. Well-designed systems can perform strongly.
  • Human scoring is not automatically superior. Humans also show severity differences, drift and bias.
  • Perfect human agreement is not required before automation. But the reference standard must be good enough for the intended use.
  • One fairness study does not certify every model. Fairness is system-, task- and population-specific.
  • Fast feedback is not the same as valid feedback. Speed is an operational advantage, not evidence of construct validity.

How Do We Know?

ETS’s 2025 chapter Automated Scoring of Open-Ended Written Responses: Possibilities and Challenges reviews the validity, fairness and technical issues involved in automated scoring for large-scale assessment.

A 2025 study in the International Journal of Artificial Intelligence in Education, Automated Scoring of Constructed Response Items in Math Assessment Using Large Language Models, reported human-like agreement on nine of ten NAEP mathematics items in a held-out set, illustrating the real potential of modern approaches.

Fairness remains a separate question. The 2025 open-access study Algorithmic Fairness in Automatic Short Answer Scoring found a minor language-background bias in the best-performing system despite strong overall scoring accuracy. A separate 2025 analysis, Is GPT-4 fair? An empirical analysis in automatic short answer grading, found encouraging consistency in the populations studied but explicitly recommended further research.

Research in Assessing Writing in 2025, Using ChatGPT to score essays and short-form constructed responses, likewise concluded that agreement can approach human scoring on some datasets while remaining inconsistent enough that high-stakes replacement still requires caution and stronger validation.

The evidence boundary is crucial: automated-scoring performance changes with task, model, prompt, language, rubric, training data and deployment design. There is no universal “AI marker accuracy” that can be transferred safely from one context to another.

For Parents and Students: Ask What Happens When the Machine Is Unsure

If automated scoring affects a consequential result, useful questions include: Is the score reviewed by a human? What evidence validates this system for this task? Can unusual responses be escalated? Is there an appeal process? Is the automated score one signal or the final decision?

The point is not to fear machines. It is to hold every measurement system—human or automated—to the same question: what does this evidence justify believing?

Bolt Direction Graph

Student response → automated score → compare with calibrated reference → inspect disagreement → test edge cases and subgroup fairness → validate across new tasks → monitor deployment → human escalation where needed → recalibrate allowed use.

Useful neighbours: Bolt Measurement Note 03 — When Two Good Teachers Give Different Marks, Bolt Measurement Note 19 — The Marker Can Drift Even When the Rubric Does Not, and Bolt Measurement Note 21 — A Test Question Can Behave Differently for Comparable Groups.