The Tutor Handbook · Volume 0136 · Series ID THB-0136
A learner scores 72 per cent on one practice set and 68 per cent on the next. The number appears to say performance fell. Yet the second set may have contained harder questions, fewer scaffolds, less familiar representations, or an adaptive platform that deliberately routed the learner towards more demanding items after earlier success. A lower raw percentage can coexist with stronger capability. The reverse is also possible: a higher score can be produced by easier work.
The tutoring problem is not to distrust every percentage. It is to decide when two scores represent sufficiently similar opportunities that a direct comparison is useful, and when changing difficulty has altered the meaning of the number.
Return to The Tutor Handbook complete series index.
The direct answer
When question difficulty changes materially, a tutor should not compare raw percentages as though the tasks stayed equal. First establish what changed in the demand: prerequisite knowledge, number of steps, representation, language load, scaffolding, novelty, time pressure, response mode, or the selection rule that chose the item. Then interpret performance against that demand. Progress is stronger when the learner can succeed on increasingly demanding work with stable or reduced support, not merely when the visible percentage rises.
This does not require a psychometric model for a three-student tutorial. It requires disciplined comparability.
A tutor can use three complementary views. The first is performance on a small stable anchor set that remains sufficiently similar across occasions. The second is performance on appropriately challenging new work. The third is qualitative evidence about how much support, time and prompting the learner needed. Together they are more informative than a single percentage from a moving target.
Why adaptive difficulty creates a measurement trap
Adaptive systems are designed to change the work. That is their point. A learner who answers well may be routed to harder items; a learner who struggles may receive easier ones. Even outside software, tutors adapt continuously. We choose a harder example when a learner is ready, reduce scaffolding, change the context, introduce a less familiar representation, or ask for explanation rather than selection from options.
This is good teaching. It becomes bad measurement only when the changed task is forgotten.
OECD technical documentation for PISA 2025 illustrates the principle at a formal assessment scale. In adaptive reading, students can be routed to different stages based on earlier performance, and the OECD percentage-correct analyses use inverse routing-probability weights because raw item exposure is not the same for everyone. PISA is a large-scale assessment with formal models and cannot be imported into private tuition. But it makes one point clearly: once item selection depends on prior performance, a simple raw percentage may no longer mean what it meant in a fixed test.
A tutor needs the practical version of that insight. Ask: did the learner receive the same kind of opportunity this time?
Difficulty is not one thing
Tutors often describe a question as “harder” as though difficulty were a single property. In practice, several demands can change at once.
A mathematics problem can become harder because it has more steps, because the structure is less obvious, because irrelevant information is added, because numbers are less friendly, because the representation changes from a diagram to prose, or because the learner must choose a method rather than imitate one.
A comprehension question can become harder because the answer is less locally stated, because several sentences must be integrated, because the wording is unfamiliar, because evidence must be selected rather than recognised, or because the response must be explained precisely.
A science item can become harder because the concept is more advanced, because two variables interact, because the data are noisy, because the learner must distinguish observation from inference, or because the question asks for transfer into a new situation.
These are different difficulty mechanisms. A learner may improve on one and not another.
That means the tutor should record the dominant change, not merely “Level 3” or “hard set”. A compact note such as “same concept; less scaffolding; unfamiliar representation” is often enough to make later interpretation more honest.
Composite case: the apparent decline
The following case is fictional and constructed for illustration.
Alicia completes a short algebra set in Week 1 and gets 9 of 10 correct. In Week 3 she gets 7 of 10. A parent sees 90 per cent followed by 70 per cent and asks why tuition is making her worse.
The tutor compares the tasks. Week 1 contained isolated one-step equations with the unknown already visible in the same position. Week 3 mixed equations with brackets, variables on both sides and worded situations in which Alicia had to construct the equation herself. The second set also removed the worked example that had appeared above Week 1.
The score fell. The demand rose.
That does not prove Alicia improved. A harder set cannot automatically be declared a success merely because it is harder. The tutor therefore uses two checks. First, a small anchor set of the Week 1 demand is repeated with different numbers; Alicia completes it accurately without the old example. Second, a delayed changed-condition item from the Week 3 family is attempted a few days later without prompting.
Now the evidence is interpretable. The older capability has been retained, while the newer, more demanding capability is emerging but not yet stable. “70 per cent” becomes one part of a developmental story rather than a verdict.
The anchor-set principle
An anchor is a small piece of work deliberately kept comparable enough to help interpret change over time.
It need not be identical. Repeating the same questions creates memory effects. A useful anchor preserves the underlying demand while changing surface details. The tutor might keep the same command type, number of reasoning steps and level of support while changing numbers, names, contexts or item order.
Anchors solve one specific problem: they provide a relatively stable reference while other work becomes more challenging.
They should remain small. If every lesson contains a large fixed test, tuition turns into continuous measurement and steals time from learning. A few representative items can be enough to answer the narrow question, “Can the learner still do the capability we had already established?”
AERO’s current Monitor progress practice guide emphasises collecting information about what students know, understand and can apply, then using it to adjust instruction. It does not prescribe a private-tuition anchor protocol. The anchor idea here is a practical implementation of the comparability problem, not a validated instrument.
Do not let the anchor become the curriculum
A stable anchor has a danger: it can become too familiar.
If Alicia sees the same structure every week, the tutor may end up measuring recognition of the anchor rather than flexible knowledge. The learner learns the measurement device.
For that reason, anchor items should be refreshed while preserving the intended demand. The tutor should also pair them with transfer items in changed contexts.
The question is not “Can the learner repeat our favourite check?” It is “Does the capability survive when the surface changes?”
A robust progress picture therefore includes both stability and transfer. Stability asks whether previously demonstrated capability remains available. Transfer asks whether it travels.
The difficulty-support pair
Difficulty cannot be interpreted separately from support.
A learner who solves a very hard problem after six tutor prompts has done something different from a learner who solves a moderately hard problem independently. Both experiences may be educationally valuable. They are not equivalent evidence of independent capability.
For each important task, the tutor should notice the pair: what demand the task carried, and what support the learner received.
Support includes more than obvious hints. It includes worked examples left visible, leading questions, peer explanations, answer options, partially completed diagrams, vocabulary supplied by the tutor, calculator access when calculation is part of the target, and repeated reminders of the next step.
Again, the aim is not bureaucratic coding. The tutor only needs enough information to avoid comparing unlike performances.
“Harder task, same support” means something different from “harder task, much more support”. “Same task, less support” can itself be evidence of progress.
Access support must not be mislabelled as easier work
A legitimate accommodation can change the conditions without changing the intended standard. Extra time, a permitted response technology, enlarged print, a reader where reading is not the construct, or another authorised access arrangement may enable the learner to demonstrate the target capability more fairly.
The tutor should not call every supported condition “easier”. That language can wrongly imply that the learner’s achievement is discounted. The right question is whether the support changes an irrelevant access barrier or performs part of the target skill.
If the support is legitimate and stable, compare like with like. If it changes between progress checks, record the change because it can affect interpretation. A learner who moves from a heavily tutor-mediated version of a support to independent use of an authorised tool may show meaningful progress even when the academic task is unchanged.
Conversely, removing a legitimate accommodation to create a “pure” test can make the evidence less valid. Independence does not mean deprivation of access.
This distinction is especially important when adaptive software automatically changes presentation, hints or navigation. The tutor needs to know which adjustments are access features and which are instructional scaffolds. If the platform does not make that distinction visible, avoid overclaiming from its score.
The practical rule is simple: preserve legitimate access; track answer-giving or task-reducing support separately. Difficulty should describe the academic demand, not the learner’s right to access the task.
Composite case: the adaptive app
The following case is fictional and constructed for illustration.
Beatrice uses an online practice tool for vocabulary. On Monday she scores 82 per cent. On Thursday she scores 74 per cent. The platform has been selecting new items based partly on previous responses, but the tutor does not know its proprietary difficulty scale.
The wrong response is to infer decline from 82 to 74. Another wrong response is to assume the app’s adaptivity guarantees progress.
The tutor separates what is known from what is not. Known: the platform changes items; Thursday contained more unfamiliar words and fewer repetitions from the previous session; Beatrice still answered most items correctly. Unknown: the precise routing rule, whether item difficulty is calibrated, and whether two displayed percentages are designed for direct comparison.
The tutor therefore uses the app as practice evidence, not as the sole progress measure. A short tutor-designed sample checks whether Beatrice can explain meanings and use selected vocabulary in novel sentences without the app’s prompts. The app remains useful, but the programme no longer treats its percentage as a universal scale.
This is an important AI and tool-use principle. A number produced by an adaptive system is not automatically a comparable measurement merely because it looks precise.
Raw scores, scaled scores and tutor judgement
Formal assessment systems sometimes use scaled scores or item-response models to place performance on a common scale when different learners encounter different items. OECD’s PISA reporting is one example. The construction of those scales requires data, calibration, assumptions and technical procedures far beyond a small tuition centre.
Tutors should not imitate the appearance of psychometrics without the substance.
Do not invent a home-made “difficulty coefficient” and multiply it by percentage correct. Do not label a ten-item worksheet “Level 4.3” and treat the number as validated. Do not average unlike tasks into a precision-looking score.
A tidy formula is not a measurement system.
Tutor judgement is still necessary, but it should be inspectable. Instead of claiming “difficulty-adjusted score = 81.6”, write what changed: “percentage lower because two items required independent model construction; anchor skill retained; new demand partly successful; no hints on four of six attempts.”
That is less glamorous and more useful.
When percent correct is still useful
The adaptive-difficulty check does not abolish percentages.
Raw percentage is useful when the item set is sufficiently similar, the scoring rule is stable, the learner had comparable time and access, and support conditions have not materially changed. It can also be useful within one session to describe exactly what happened on that set.
The mistake is not calculating 7 out of 10. The mistake is treating every 7 out of 10 as the same achievement.
Tutors should ask what comparison the number is being asked to support. If the question is “How many did she answer correctly today?”, the raw percentage answers it. If the question is “Has her algebra capability improved across six weeks of increasingly difficult work?”, the percentage alone does not.
Different questions require different evidence.
The challenge ladder
One practical way to organise adaptive work is a challenge ladder.
The ladder does not need numerical levels. It can describe what changes from one rung to the next.
For example, in comprehension, the tutor may move from locating explicitly stated information, to connecting information across sentences, to inferring a relationship from evidence, to justifying the inference with precise textual support, and finally to transferring the same reasoning to a less familiar text.
In mathematics, the learner may move from executing a known procedure, to choosing between plausible procedures, to constructing a representation, to solving in a changed context, and finally to explaining why the method works or why an alternative fails.
The ladder helps the tutor say what “harder” means. It also prevents challenge inflation, where a tutor merely uses longer questions and calls them more advanced.
A challenge ladder is a teaching map, not a validated scale. Its value is transparency.
AERO’s Extend and challenge guidance advises extending students beyond their current level of mastery through appropriate challenge while avoiding unnecessary cognitive overload. That supports the general instructional direction, not any particular ladder used in tuition.
The floor and ceiling problem
A task can be too easy to show growth.
If a learner scores 10 out of 10 on a very simple set in Week 1 and 10 out of 10 again in Week 6, the score has no room to reveal improvement beyond the set’s ceiling. The learner may now be capable of much more, but the instrument cannot show it.
A task can also be too hard. Repeated near-zero performance tells the tutor little about fine-grained growth because the learner has too few successful opportunities to reveal partial change.
This is why the Tutor Handbook Measurement Range Check matters. It owns the broader question of whether a progress check sits in a useful range. The adaptive-difficulty check adds a different decision: once the tutor changes challenge to escape a floor or ceiling, how should the new result be interpreted relative to the old one?
The answer is to preserve a bridge. Keep a small amount of overlapping demand while introducing the new range.
The bridge prevents a false discontinuity.
Changed representations are a difficulty change
Tutors sometimes believe two tasks are equally difficult because they target the same syllabus objective. That can be wrong.
A learner may solve a ratio problem when represented as a table but struggle when the same relationship appears in prose. Another may understand a science process in a labelled diagram but lose the structure when asked to explain it in continuous writing. A vocabulary item recognised in multiple choice may not be retrievable in free production.
The underlying knowledge target may be related, but the retrieval and representation demands differ.
When the tutor changes representation intentionally, record it. If performance falls, investigate whether the difficulty lies in the concept or in coordinating the new representation.
Then test both. Return briefly to the established representation to confirm the old capability remains, and provide a fresh opportunity in the new one after instruction.
This prevents “same topic” from being mistaken for “same task”.
Adaptive difficulty in a three-learner group
Small-group tuition creates another comparability problem: three learners may receive different questions at the same time.
That can be good orchestration. One learner needs a repair item, another an alignment item, and another a frontier challenge. Equal worksheets are not automatically fair or educationally sensible.
But the tutor should not rank the learners using raw scores from different work.
If Emily scores 9 out of 10 on foundation items while Faith scores 6 out of 10 on advanced transfer tasks, “Emily is doing better” is not a defensible conclusion from those numbers.
The tutor can compare each learner with their own defined targets and use common anchor items when a shared comparison is genuinely needed. The group can still discuss methods together without turning every result into a competition.
This matters for parent communication as well. Reports should make the task demand visible enough that percentages are not stripped from context.
A calm sentence such as “Faith’s raw accuracy was lower because we moved into unscaffolded transfer; her prior anchor skill remained secure” communicates more truth than a colour-coded leaderboard.
Composite case: the tutor accidentally makes progress invisible
The following case is fictional and constructed for illustration.
Ciara begins a term needing heavy scaffolding to write analytical paragraphs. Early practice uses sentence frames and visible success criteria. Her paragraphs are mostly complete.
Six weeks later, the tutor removes the sentence frame, gives a less familiar text and asks Ciara to choose evidence independently. Her first attempt is messier and contains more errors.
If the tutor tracks only surface correctness, progress appears to reverse.
But the task has changed in three important ways: support is lower, text novelty is higher, and decision-making demand is greater.
The tutor preserves an old-demand anchor by asking Ciara to write one short paragraph with the familiar scaffold. She does so fluently. Then the tutor returns to the unscaffolded task and observes where independence breaks.
The conclusion is not “she got worse” or “she definitely improved”. It is more precise: supported performance has become stable; the route is now being tested under greater independence; evidence selection is the current weak link.
That conclusion directly informs teaching.
The danger of teaching to the metric
Once tutors start tracking progress, they can unconsciously optimise the number.
If parents respond positively to rising percentages, the tutor may keep work easy. If a centre celebrates high completion rates, difficult open-ended tasks may disappear. If software rewards streaks, learners may choose familiar items. A measure intended to reveal learning begins to shape the environment in ways that inflate itself.
This is a form of metric capture.
The antidote is to keep the educational job primary. The tutor chooses challenge because it is the right next learning demand, not because it protects the graph.
A useful progress system should tolerate temporary mess when the learner moves into a genuinely harder task. It should also detect when “harder work” is merely chaos, overload or poor sequencing.
The tutor needs both ambition and calibration.
Parent communication: show the moving target
Parents often receive percentages without seeing the work.
When difficulty is changing, reports should include one sentence describing the change in demand.
For example: “Accuracy moved from 84% to 76%, but the second sample removed the worked example and required method selection.” Or: “Raw score stayed around 70%, while question type moved from direct recall to explanation using unfamiliar data.” Or: “Accuracy rose, but support also increased this week, so we will verify independently next session.”
These statements do not excuse low performance. They prevent overclaiming.
They also help parents avoid an unhelpful pattern in which every lower score triggers more worksheets at the old level. Sometimes the correct next move is to keep the learner in the harder range long enough to learn from it.
Parents should still ask whether challenge is productive. A tutor who constantly says “the work is harder” without preserving any stable evidence can hide stagnation. That is why anchors matter.
A practical adaptive-difficulty record
A short record can contain five fields: target capability; dominant change in demand; support available; performance observed; next verification.
Example: “Target: infer cause from data. Change: unfamiliar graph plus distractor variable. Support: no hints, vocabulary clarified only. Performance: correct conclusion, weak justification. Next: changed graph after three days.”
This is not meant to become paperwork for every question. Use it for transitions where raw score would otherwise mislead.
The record forces the tutor to state why the task was harder rather than merely asserting it.
It also makes handover easier. A future tutor can see whether a lower score represented regression, frontier work, reduced support or an access issue.
When to hold difficulty steady
Adaptation is not always better.
If the tutor changes difficulty after every response, the learner never receives enough repeated exposure to stabilise a method. The tutor may also lose the ability to tell whether instruction worked because the conditions keep moving.
Sometimes the correct decision is to hold the demand steady for several attempts.
This is especially important after introducing a new representation, after removing a scaffold, or when the learner is practising a complex routine. Let the learner accumulate enough experience for the tutor to observe change.
Then alter one meaningful dimension.
The Tutor Handbook Stability Window owns the broader logic of holding a new route steady long enough to learn from it. Here the application is measurement: stable windows create interpretable evidence before the next difficulty change.
Failure modes
The first failure mode is raw-percentage literalism: 80 is assumed better than 70 without checking the work.
The second is difficulty theatre. The tutor labels work “advanced” but cannot say what changed in the cognitive demand.
The third is challenge as excuse. Every weak result is defended as “harder work”, so no unfavourable evidence can ever count.
The fourth is false psychometrics. Home-made scales and weighted formulas create numerical precision unsupported by calibration.
The fifth is anchor overtraining. Stable checks become so familiar that they measure memory of the check.
The sixth is difficulty confounding. The tutor changes content, representation, timing, support and response mode simultaneously, then cannot tell which change mattered.
The seventh is support amnesia. A hard task solved with extensive prompting is recorded as if it were independent performance.
The eighth is ceiling maintenance. The tutor keeps easy work because rising percentages look reassuring.
The ninth is frontier addiction. The tutor constantly escalates challenge and never lets competence stabilise.
The tenth is group ranking across unlike tasks.
Each failure comes from forgetting that a score belongs to a performance under conditions.
Research limits
Formal adaptive testing uses calibrated item pools, statistical models and routing rules. OECD PISA documentation describes these processes at a scale and level of technical control that ordinary tuition does not possess. It would be misleading to imply that a small tutor-designed challenge ladder is equivalent.
AERO’s monitor-progress and extend-and-challenge practice guides synthesise research-informed school guidance, not private-tuition trials. Their principles can inform instructional judgement, but they do not validate a specific eduKate three-student protocol.
The adaptive-difficulty check in this volume is therefore a professional reasoning framework. It is not a standardised assessment, a psychometric conversion formula, or proof that any particular learner has improved.
Its purpose is narrower: prevent the tutor from comparing scores that no longer represent the same opportunity.
Source map and evidence status
Current large-scale assessment methodology: OECD, PISA 2025 Results (Volume I), Technical notes on analyses in this volume. The technical notes describe multistage adaptive routing and the use of inverse routing-probability weights for certain percentage-correct analyses.
Assessment-scale background: OECD, PISA 2022 Results (Volume I), The construction of reporting scales and of indices from the student context questionnaire. This describes placing item difficulty and student proficiency on common scales in formal assessment.
Research-informed instructional guidance: Australian Education Research Organisation, Monitor progress, updated 14 May 2026.
Research-informed challenge guidance: Australian Education Research Organisation, Extend and challenge, updated 14 May 2026.
These sources support the principles of explicit demand, monitoring and appropriate challenge. None should be read as experimental validation of this private-tuition framework.
A delayed comparability check
When a learner first moves into a harder range, do not decide too quickly whether the transition succeeded.
Preserve a small anchor, teach into the new demand, then check again after enough time for adaptation. If old capability remains secure and performance on the harder demand improves with stable or reduced support, the progress claim becomes stronger.
If the anchor deteriorates while frontier work expands, investigate whether the new route is displacing foundations. If the harder work never becomes more independent, the tutor may be maintaining performance through support rather than building capability.
Changed-condition checks matter too. Can the learner carry the skill into a different representation, different context or different item order without returning to the tutor for the route?
The answer should update the learning plan.
Final return
Progress is not a race to the highest percentage.
A learner can grow while the score falls, stagnate while the score rises, or appear stable while the task quietly changes underneath.
The tutor’s responsibility is to keep the change visible.
Ask what became harder. Ask what support remained. Preserve a small comparable anchor. Use fresh transfer tasks. Refuse false precision. Do not rank learners across unlike work. Explain changing demand to parents. Hold difficulty steady when interpretation requires it, and raise it when the learner is ready.
A percentage describes a performance.
Progress describes a change in capability.
Those are related, but they are not the same thing.