Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

The Tutor Handbook Vol No.0125 | The Criterion Drift Check — How a Tutor Keeps the Standard Stable Enough to Read Progress Without Quietly Becoming Harsher, Softer or More Demanding Over Time

The Tutor Handbook · Volume 0125 · Series ID THB-0125

Return to The Tutor Handbook

A tutor marks a learner’s paragraph in September and calls the evidence “adequate but not yet secure”.

Three weeks later, the learner produces a paragraph of roughly the same quality. The tutor now calls it “weak”.

Nothing obvious in the learner’s work explains the change.

The tutor has simply seen better writing in the meantime. Their internal standard has moved.

Or the opposite happens. After weeks of seeing a learner struggle, an answer that would once have been judged incomplete begins to feel impressive. The tutor has become more generous because the work is better than the learner’s recent history, even though it has not yet reached the declared target.

Both errors matter because tutoring depends on comparing evidence across time. If the standard moves silently while the learner is changing, the tutor can no longer tell which movement belongs to the learner and which belongs to the judge.

This article owns one decision: how does a tutor keep the criterion stable enough to read progress while still allowing the target to change legitimately when the educational job changes?

The direct answer

A tutor should separate three things that are often blended together:

  • the criterion — what counts as an adequate performance for the present target;
  • the evidence — what the learner actually did under known conditions; and
  • the target state — the level or kind of performance the tutor is currently trying to build.

The criterion should not become harsher or softer merely because the tutor is tired, impressed, disappointed, familiar with the learner, comparing with a stronger peer, or remembering a recent run of unusually good or weak work.

But a criterion is not frozen forever. It may legitimately change when the task changes, the curriculum demand changes, the learner moves from supported practice to independent performance, or the tutor deliberately raises the target after the earlier target has been met.

The discipline is therefore not “never change the standard”. It is:

Do not let the standard change invisibly.

When a target changes, name the change. When the target is supposed to be stable, use anchors, explicit criteria and occasional re-checks to see whether the judgement itself has drifted.

What criterion drift is

Criterion drift occurs when the meaning of “good enough”, “secure”, “accurate”, “well explained” or another judgement category changes over time even though the intended target has not.

The drift can be stricter.

A tutor sees increasingly sophisticated work and begins expecting more detail than the original task requires. Earlier acceptable answers would now be rejected.

The drift can be looser.

A tutor becomes accustomed to persistent difficulty and gradually rewards partial evidence as though it met the target. Earlier inadequate answers would now be accepted.

The drift can also be sideways.

The tutor starts valuing a different feature. An English response originally judged mainly for evidence selection becomes increasingly judged for elegance. A Mathematics solution originally checked for method validity becomes increasingly judged for brevity. A Science explanation originally judged for causal accuracy becomes increasingly judged for technical vocabulary.

Those may all be useful qualities. The problem is not that they matter. The problem is that the tutor has changed what the judgement means without recording that the target has changed.

Criterion drift is not the same as learner progress

This distinction sounds obvious until real work arrives.

Suppose Beatrice’s first three responses contain relevant evidence but weak explanation. The tutor says, “You are selecting the right evidence; now we need to make the link explicit.”

Over several sessions, Beatrice becomes better at linking evidence to the claim.

At that point, the tutor reasonably begins paying more attention to qualification, counterargument and precision.

That is not necessarily criterion drift. It may be a deliberate progression to a more demanding target.

Now suppose the tutor looks back at Beatrice’s first response and says, “This was actually very poor because it had no qualification.”

That retrospective judgement can be misleading if qualification was not part of the target being assessed at that time.

A learner can move through changing targets. The tutor still needs an honest record of what each earlier judgement meant.

Progress is difficult to interpret when yesterday’s criterion is rewritten using today’s expectations.

Why the problem appears in tutoring

Tutoring is unusually vulnerable to criterion drift because the tutor knows the learner well.

That familiarity is valuable. It allows the tutor to notice hesitation, strategy choice, recovery, confidence, independence and patterns that a one-off assessor may miss.

Familiarity can also contaminate judgement.

The tutor knows that Alicia has worked hard.

The tutor remembers that Denise used to freeze on this kind of task.

The tutor has seen Emily produce much stronger work than the current answer.

The tutor knows Faith was tired after school.

All of that context can matter for instruction. It does not automatically belong inside the judgement of whether this particular performance met a declared criterion.

A humane tutor does not need to become a blind examiner. The tutor needs to know when context should change the next teaching decision and when it should not change the meaning of the evidence just collected.

The contrast effect

A response does not arrive in isolation. It arrives after something else.

If the tutor has just read an exceptional response, an adequate response may feel weak.

If the tutor has just read two confused responses, a merely competent answer may feel excellent.

In a three-learner tutorial this can happen within minutes.

Alicia gives a concise, rigorous explanation. Beatrice follows with a longer but still valid explanation. If the tutor has unconsciously recalibrated to Alicia, Beatrice may be judged against Alicia rather than against the task.

This is one reason the existing Peer Baseline Trap matters. The fastest or strongest learner should not become the hidden criterion for everyone else.

The same principle applies across time. Yesterday’s unusually weak or strong batch of work should not silently reset today’s standard.

The familiarity effect can move in both directions

Tutors sometimes become more lenient with learners they know well.

They understand what the learner “meant”. They can reconstruct missing steps. They remember the discussion that preceded the written answer. They know that the learner usually handles the concept correctly.

That extra knowledge can turn an incomplete artefact into a complete story in the tutor’s mind.

The opposite also happens.

A tutor who knows a learner can do better may become harsher than the task warrants. A perfectly adequate response is treated as weak because it is below the learner’s personal best.

Instructionally, “you can do better than this” may be useful.

Measurement-wise, it is different from “this does not meet the criterion”.

Those statements should not be collapsed.

The criterion anchor

A criterion anchor is a compact reminder of what the present judgement is meant to recognise.

It does not need to be an enormous rubric.

For a particular task it may be as small as:

  • selects evidence that directly bears on the claim;
  • explains the connection rather than merely quoting;
  • keeps the claim within what the evidence supports.

For a worked Mathematics item:

  • method is mathematically valid;
  • necessary reasoning is visible where the task requires it;
  • final result is consistent with the working.

For an oral explanation:

  • identifies the relevant idea;
  • gives a coherent causal or logical account;
  • responds to one follow-up without the tutor supplying the core reasoning.

The anchor should be specific enough to constrain judgement and small enough that the tutor can actually use it.

A forty-line rubric that is never consulted is not a stable criterion. It is decoration.

Anchors should not become answer keys

An anchor is for judgement. It is not automatically something to reveal in full before every attempt.

If the learner’s target includes independently deciding what matters, an over-detailed checklist can become a route map that supplies the thinking.

The existing Criteria Disclosure Gate owns that disclosure problem.

For criterion stability, the tutor may need a precise private or shared anchor even when the learner receives a simpler goal statement.

The important point is that the tutor knows what is being judged and does not improvise a new standard after seeing the response.

A fictional composite case: when “better” becomes “good enough”

This is a fictional composite tutoring case. It does not describe a real customer or learner.

Ciara is practising short written explanations. Early responses omit the link between evidence and conclusion.

The tutor sets a narrow target: each response must contain the relevant evidence and an explicit sentence explaining why that evidence supports the conclusion.

Week one: Ciara supplies the evidence but no link.

Week two: she supplies the evidence and a vague link.

Week three: she supplies a clear link.

Because the improvement is obvious, the tutor writes, “Excellent — secure.”

In week four, the tutor asks for a new response under the same stated criterion. Ciara again supplies evidence and a clear link, but the wording is plain and the conclusion lacks nuance.

The tutor writes, “Needs improvement.”

Why?

The tutor has started expecting sophistication that was not part of the original criterion.

There are two responsible routes.

One is to keep the original criterion for this progress check and judge the work as meeting it, while separately noting that the learner is ready for a new target.

The other is to announce that the target has changed: the next phase now adds qualification or precision.

What is not responsible is to declare the earlier standard inadequate only after the learner reaches it.

Stable enough does not mean identical task after identical task

A stable criterion can be applied to varied evidence.

In fact, variation is often necessary if the tutor wants to know whether the capability transfers beyond one rehearsed form.

The task surface can change while the judgement rule remains stable.

The learner might explain the same reasoning in a new context, use a different data display, solve an unfamiliar problem with the same underlying structure, or write about a different text.

This helps separate criterion stability from item repetition.

The goal is not to make every check identical. It is to preserve the meaning of the judgement while allowing the evidence to become fresher and more varied.

Stable criteria and changing conditions

Sometimes the conditions legitimately change.

The tutor may remove a scaffold.

The time limit may become realistic.

The learner may move from guided to independent work.

The task may include more competing information.

Now the tutor is no longer collecting evidence under the same conditions.

That does not make the comparison useless. It makes the comparison conditional.

Instead of saying, “The learner went from 7/10 to 6/10, so learning declined,” the tutor can say, “The later check was completed with less support, so the raw score is not directly equivalent. The learner maintained most of the performance under a harder independence condition.”

The existing Learning Claim owns the discipline of keeping claims inside the evidence.

Criterion stability helps because it prevents a second source of change — the judge — from being mixed into already-changing task conditions.

The calibration pause

A tutor does not need a formal moderation meeting before every worksheet.

A short calibration pause is often enough.

Before marking a new set, ask:

What would have counted as meeting this target last week?

Then inspect the anchor or one de-identified exemplar if one exists.

This simple step is especially valuable when:

  • the tutor has just seen unusually strong or weak work;
  • the learner has improved quickly;
  • several tutors share a programme;
  • the judgement affects a route change;
  • the target contains qualitative features such as explanation, organisation or reasoning.

The aim is not bureaucratic consistency for its own sake. It is to keep the educational decision readable.

Re-reading an old sample can expose drift

Occasionally, a tutor can re-read an earlier anonymised or de-identified sample without looking first at the old judgement.

The question is not, “Can I reproduce exactly the same mark?”

Human judgement contains variation.

The useful question is whether the meaning has moved enough to change the educational decision.

If a response previously judged secure would now be called weak under the same stated target, inspect why.

Maybe the earlier judgement was too generous.

Maybe the later judgement is too harsh.

Maybe the target actually changed and the documentation failed to show it.

Any of those discoveries can improve the system.

Do not turn this into constant re-marking. The purpose is a periodic drift check where the consequence of changing standards would matter.

More than one tutor: calibration becomes a handoff problem

A learner may move between tutors, receive cover tuition, or have work reviewed by a tutor coach.

Now criterion stability is no longer just within one person’s head.

Two tutors can use the same words and mean different things by “secure”.

The solution is not to force perfect agreement on every borderline response.

Instead:

  1. define the target;
  2. inspect the same small evidence sample;
  3. state what features each tutor used;
  4. identify where disagreement changes the next decision;
  5. if necessary, collect a fresh sample that better separates the competing interpretations.

The existing Evidence Conference owns disagreement between tutors. A criterion drift check asks an earlier question: are they even applying the same declared standard?

Do not solve disagreement by averaging judgements

One tutor says “secure”. Another says “not secure”. Calling the answer “half secure” does not resolve the meaning.

The tutors may be looking at different features.

One may care about the final answer. The other may care about the reasoning.

One may assume the scaffold is allowed. The other may be judging independence.

One may be using the current target. The other may have silently advanced to the next target.

Expose the criterion before trying to reconcile the score.

Averaging can hide the very disagreement that needs investigation.

Rater research offers a warning, not a tutoring formula

Formal assessment research has long examined variation in human ratings, including differences in severity and the possibility that rating behaviour changes over time.

That literature is relevant because it shows that human judgement is not automatically constant merely because the same person or rubric is involved.

But the transfer to small-group tuition must be careful.

Operational examinations, professional assessments and peer-rating studies have different stakes, training, numbers of raters and scoring systems. They do not validate a particular “drift index” for a tutor working with three learners.

For example, Leckie and Baird examined rater effects in an English curriculum assessment in England and found complex patterns in severity across rating periods. Other studies have examined rater severity changes in professional or peer-assessment settings. These bodies of work justify vigilance about rating stability. They do not prove that a given tutor’s weekly judgement will drift, nor do they supply a universal acceptable threshold.

The practical framework here is therefore research-informed governance, not a psychometric instrument.

The tutor should distinguish three changes

Whenever a judgement shifts, ask which of these changed:

1. The learner changed

The same target, comparable conditions and stable criterion now produce stronger evidence.

That supports a progress claim.

2. The task or condition changed

The learner may be working with less support, more complexity, less time or a different form.

The result may still be valuable, but direct comparison needs qualification.

3. The judgement standard changed

The tutor now expects more, less or something different.

That may be legitimate if deliberate. If accidental, repair the comparison before changing the learner’s route.

This three-change check prevents the tutor from interpreting every difference as learner change.

Raising the bar responsibly

There is a point at which keeping the old target becomes educationally wasteful.

The learner has met it repeatedly, under delay, variation and appropriate independence.

The tutor should move on.

A responsible target raise has a visible boundary.

For example:

“Up to this point, success meant selecting relevant evidence and explaining the link. That is now stable enough to move forward. From today, the target also includes qualifying the claim when the evidence has limits.”

Now earlier work remains valid evidence about the earlier capability. Later work is judged against a stronger target.

The learning record becomes a sequence rather than a rewritten history.

Lowering the bar can also be legitimate

A task can reveal that the present target is badly sequenced.

The tutor may discover a prerequisite weakness, an access barrier or excessive simultaneous demand.

The responsible response may be to narrow the target temporarily.

That is not the same as becoming lenient.

“Today I am only judging whether the learner can identify the relevant evidence; the explanation link will be rebuilt next.”

The criterion is actually more precise, not softer.

Difficulty and standard are not synonyms.

A narrower target can be demanding within its own scope.

Criterion drift in feedback

Feedback itself can reveal movement in the standard.

Week one, the tutor says, “Make the link explicit.”

Week two, “Good; now make the link precise.”

Week three, “Avoid repetition.”

Week four, “Use a more sophisticated connective.”

There may be sensible progression here.

But if the learner keeps being told “not yet” because a new requirement appears each time the previous one is met, the learner experiences a moving finish line.

That can damage agency and make success criteria feel arbitrary.

A tutor should be able to tell the learner, “You met the target we were working on. We are now opening the next one.”

Progress deserves an honest receipt before the route becomes harder.

Criterion drift in oral tutoring

Written work leaves an artefact. Oral answers are easier to reinterpret.

The tutor hears tone, hesitation and fluency. They may remember only the final polished answer after several prompts.

A simple record helps:

  • first response;
  • level of support;
  • criterion being judged;
  • final status.

This need not become a transcript.

The point is to preserve enough evidence that “secure” next week means roughly what “secure” meant this week.

The Recording Window is relevant here: documentation should capture decision-relevant evidence without causing the tutor to miss live reasoning.

Three learners should not create three private standards

In a small group, individualisation does not mean that the same declared target should have arbitrary meanings for different learners.

Two learners may need different supports.

One may have a legitimate access accommodation.

One may be working on a different stage of the pathway.

Those differences should be visible.

But if both learners are being judged on the same independent target under comparable conditions, the tutor should not quietly accept less from one because they expect less of that learner.

That turns history into the standard.

Likewise, a learner who usually excels should not be required to exceed the criterion merely to receive the same acknowledgement of success.

Personal bests are useful for coaching. Criteria are useful for deciding whether a target has been met. They are related but not identical.

Access supports are not criterion corruption

A stable criterion does not require removing legitimate access support.

If a learner has an accommodation that enables access without performing the target skill, keep it stable and visible.

The Accommodation Baseline owns the broader problem.

Criterion stability means judging the target capability consistently under the intended access condition.

Do not treat an accommodation as cheating.

Do not quietly add or remove it between checks and then attribute the score movement entirely to the learner.

A parent-facing explanation

Parents often see marks and comments rather than the criterion history.

Suppose a learner receives “secure” one month and “developing” the next.

A parent may reasonably ask whether performance declined.

The tutor should be able to distinguish:

“The learner’s performance declined under the same target.”

from:

“We raised the target after the earlier one became secure.”

from:

“The later task required more independence, so the labels are not directly equivalent.”

That is more informative than defending the latest judgement as obviously correct.

It also protects the learner from a narrative in which every harder target erases the progress that justified moving to it.

Do not manufacture precision

A tutor may be tempted to score criterion stability numerically.

“Rater severity index: 0.73.”

Unless the measure has a justified model, reliable data and an interpretable scale, the number is cosmetic.

A practical tuition programme often needs much less:

  • target statement;
  • criterion anchor;
  • conditions;
  • short exemplar or reference point where useful;
  • current judgement;
  • note when the target changes.

This is enough to make many important decisions auditable without pretending the programme is running a formal psychometric study.

AI-assisted marking can drift too

AI can generate feedback or propose scores, but its apparent consistency should not be assumed.

Outputs can change with prompts, model versions, surrounding context or small differences in wording. A model can also apply an inappropriate criterion confidently.

If AI assists judgement, the tutor still owns the criterion.

Keep the human-readable target visible. Check a sample of outputs. Record material changes in the tool or prompt if comparisons depend on them. Do not compare a learner’s September score from one system with a November score from a changed system and call the difference pure learner progress.

The existing AI material guidance protects learner-facing content. Criterion drift adds another concern: the judging system itself can change.

A lightweight drift-check routine

Before a consequential progress decision:

  1. Name the target. What capability is being judged now?
  2. Retrieve the criterion anchor. What counted as adequate when this target began?
  3. Check conditions. Is support, timing, format or access materially different?
  4. Judge current evidence before comparing with the learner’s history.
  5. Compare with one or two earlier samples if needed.
  6. Ask whether a changed judgement reflects learner change, condition change or criterion change.
  7. If the target is advancing, record the boundary.
  8. If drift is suspected, recalibrate before changing the route.

This is not a demand for perfect scoring reliability.

It is a way to stop a hidden moving standard from masquerading as learner evidence.

When not to keep the old criterion

Do not preserve a criterion merely because it is old.

Change it when:

  • it no longer matches the learning objective;
  • the learner has clearly moved into a new stage;
  • official curriculum or task requirements legitimately change;
  • the earlier criterion was discovered to be invalid or badly specified;
  • the task purpose changes from supported learning to independent verification.

But make the change explicit.

A corrected criterion may require reinterpreting some earlier evidence. Say so rather than pretending the history was always clear.

Failure modes

The rising finish line. Every learner success causes a new demand to appear before the old success is acknowledged.

The sympathy discount. Persistent difficulty makes the tutor accept evidence that no longer meets the stated target.

The star penalty. A high-performing learner must exceed the criterion to be told they met it.

The peer recalibration. The strongest learner in the group becomes the standard for the others.

The history rewrite. Earlier work is rejudged using criteria that were introduced later.

The rubric theatre. A detailed rubric exists but actual judgements are made from impression.

The condition blur. Reduced support or increased task demand is mistaken for a tougher criterion, or vice versa.

The AI stability illusion. Automated scores are treated as perfectly comparable even though model, prompt or context has changed.

The false number. A homemade reliability score creates certainty the evidence does not warrant.

Research limits

Educational assessment research gives good reason to take rater effects and rating stability seriously, but much of that research concerns formal scoring systems, examinations, professional assessments or peer assessment.

That setting is not the same as a Singapore three-student tutorial.

This article therefore does not claim that tutor judgements follow the same statistical patterns or that formal rater-training methods can be imported unchanged.

The strongest direct educational guidance relevant to ordinary tutoring comes from broader formative-assessment principles: define what students should know and do, check their understanding and application, use evidence to adapt instruction, and monitor implementation rather than assuming it works.

AERO’s progress-monitoring guidance supports checking what learners understand and can apply so instruction can respond to gaps. EEF’s implementation guidance emphasises that evidence-informed approaches only matter when they are enacted coherently in day-to-day practice. These sources support the need for clear targets and readable evidence; they do not validate a particular criterion-drift protocol for private tuition.

The rater studies are therefore used as a warning about human judgement, not as proof of a tutoring effect.

Sources and further reading

The final return

A tutor’s judgement is part of the learning system.

That means the tutor is not merely reading change. The tutor can introduce change into the measurement.

The learner gets better.

The tasks change.

Support fades.

Expectations rise.

The group develops.

New curriculum demands arrive.

All of that movement is normal.

The safeguard is not to pretend that the educational world can be frozen. It is to keep track of which part moved.

When the learner changes under a stable target, say so.

When the conditions become harder, say so.

When the target advances, say so.

And when the tutor’s own standard appears to have drifted, recalibrate before converting that drift into a story about the learner.

A criterion should be stable enough to make progress visible and flexible enough to follow a legitimate change in the educational job.

The crucial word is visible.

If the standard changes, the learner should not have to discover it by repeatedly reaching a finish line that moves after they arrive.