The Tutor Handbook · Volume 0100 · Series ID THB-0100
The Tutor Handbook: Complete series index.
A learner gets a Mathematics question wrong.
The tutor looks at the topic heading and concludes that the Mathematics is weak.
That conclusion may be right. It may also be one layer too fast.
The question might require the learner to read a dense paragraph, hold three conditions in mind, interpret an unfamiliar diagram, recognise a unit conversion, choose a method, execute several operations and then express the result in a format the learner has rarely seen. A wrong final answer tells us that the whole task failed. It does not automatically tell us which requirement failed first.
The Task Purity Check asks whether the task gives a reasonably clean view of the capability the tutor intends to judge, or whether other requirements are contributing so much difficulty that the result cannot honestly be read as evidence about that capability alone.
This is not a demand for artificially easy work. Real school tasks often combine several capabilities. Learners eventually need to perform under that complexity. The professional problem is different: when the tutor is making a diagnosis, changing a learning route or claiming that one skill is weak, the tutor needs to know what the task actually required.
Quick Read
- A task can be difficult for reasons other than the target skill.
- Reading load, unfamiliar vocabulary, visual layout, memory demand, timing, interface demands, background knowledge and answer format can all alter performance.
- These demands are not automatically irrelevant. The key question is whether they belong to the capability being judged.
- A tutor should name the target construct in plain language before interpreting the result.
- Then list the additional operations the task requires.
- If a non-target demand plausibly caused the failure, change one condition and recheck.
- Do not remove legitimate complexity permanently. Use cleaner tasks for diagnosis, then return to integrated performance.
- Accessibility support should remove irrelevant barriers without supplying the target thinking.
- When two task forms give different results, investigate the difference instead of averaging the scores.
- Task purity is strongest when success and failure can be linked to the intended operation with minimal ambiguity.
- Task purity is never perfect. The goal is responsible interpretation, not laboratory isolation.
1. Start by Naming the Capability You Think You Are Measuring
Before a tutor decides that a learner is weak, the tutor should be able to complete this sentence:
I am trying to find out whether the learner can ______.
“Do this question” is not precise enough. “Understand fractions” is usually still too broad. Better descriptions name an observable operation: compare two fractions with unlike denominators; identify the relevant evidence for an inference; reconstruct a causal chain in a Science explanation; select an algebraic method when the topic is not announced; plan a paragraph that answers a specific writing purpose.
Once the target is named, task purity becomes visible. If the target is algebraic method selection, then difficult reading may be an extra requirement. If the target is solving algebra from real-world language, reading is part of the intended job. The same feature can therefore be irrelevant in one diagnostic question and essential in another.
This is why the phrase “make the question easier” can be misleading. The tutor is not necessarily reducing the intellectual standard. The tutor may be removing a requirement that does not belong to the current diagnostic question so the target capability can be seen more clearly.
2. One Task Often Contains a Stack of Hidden Operations
Consider an illustrative Mathematics problem. The learner must read a paragraph about two water tanks, infer that the relevant relationship is proportional, identify which quantity is the base, convert litres to millilitres, represent the relationship algebraically, solve the equation and present the answer in the requested unit. Calling the result “a ratio question” hides most of what the learner actually had to coordinate.
The tutor can unpack the stack into operations: decode the language; identify givens; suppress irrelevant details; choose a representation; retrieve a method; execute accurately; track units; check plausibility; express the final answer. Failure at the end could originate almost anywhere in that chain.
The same problem appears in English. A comprehension answer may depend on vocabulary, pronoun reference, inference, evidence selection, answer scope and sentence construction. A learner may understand the passage but lose the mark through answer formulation. Or the learner may write beautifully while selecting evidence that does not support the claim.
Task purity begins when the tutor stops treating the page label as the mechanism and reconstructs the actual operations the learner must perform.
3. Construct-Irrelevant Difficulty Is a Useful Assessment Idea
Educational measurement uses the idea of construct-irrelevant variance for performance differences caused by factors that are not part of the intended construct. ETS guidance gives a simple version of the issue: if a question intended to measure one capability also requires knowledge that is not part of that capability, the extra requirement can weaken the validity of the interpretation.
That principle transfers usefully to tutoring, but with an important boundary. A tutor is not administering a large-scale standardised assessment. A tutoring task is often deliberately diagnostic, instructional or mixed-purpose. The tutor therefore does not need formal psychometric validation before using a question. What the tutor does need is interpretation discipline.
Ask: if the learner fails, which parts of the task could plausibly explain the failure? If the answer includes several non-target demands, do not make a strong claim about one target capability from that task alone.
This is particularly important when a learner has legitimate accessibility needs or when the task uses unfamiliar contexts, complex interfaces or dense language. Removing an irrelevant barrier can improve the quality of the evidence without lowering the educational aim.
4. Reading Load Can Masquerade as Subject Weakness
A learner performs poorly on word problems but accurately solves the same underlying mathematical relationships when the information is presented as a diagram or short statement. That difference does not prove the Mathematics is strong. It does tell the tutor that “weak Mathematics” is too coarse a diagnosis.
The next move is not necessarily to remove language from Mathematics forever. Real examinations and real life require learners to interpret language. Instead, the tutor can separate the questions. First: can the learner perform the mathematical operation when the relationship is made visible? Second: can the learner recover that relationship from language? Third: can the learner coordinate both under realistic conditions?
Those three performances answer different educational questions. The first tests the mathematical operation more cleanly. The second tests interpretation-to-representation transfer. The third tests integrated performance.
A tutor who jumps directly from integrated failure to “reteach the whole topic” may spend weeks repairing a capability that was already available while leaving the real interface problem untouched.
5. Memory Demand Can Make Available Knowledge Look Absent
Some tasks require the learner to hold several intermediate results, conditions or instructions in mind while continuing to reason. When the target capability is not working-memory coordination itself, that extra load can obscure what the learner knows.
Imagine Ciara, a fictional learner used throughout the Tutor Handbook, explaining a Science mechanism. Orally, with the diagram visible, she produces the correct causal chain. When required to read a long prompt, remember three specified conditions and write the explanation without the diagram, she omits the second step. It would be premature to conclude that she does not understand the mechanism.
The tutor can change one condition: allow the givens to remain visible, or let Ciara annotate them before explaining. If the causal chain returns, the evidence points toward coordination or task-management load rather than absence of conceptual knowledge. The tutor can then train integration deliberately rather than reteach the mechanism from the beginning.
This is not a promise that visible supports should remain forever. It is a diagnostic move that identifies which capability needs to be built next.
6. Unfamiliar Format Can Create a False Weakness
A learner may know the content and still fail because the response format is unfamiliar. Computer interfaces, drag-and-drop tasks, tables, multi-select questions, unusual answer boxes, unfamiliar diagram conventions or new notation can introduce extra operational demands.
ETS research on technology-based assessment has examined whether interface demands create construct-irrelevant variance. The lesson for tutoring is not that digital tasks are invalid. It is that the medium can become part of the task whether the tutor intends it or not.
If the target is conceptual understanding, compare performance across a familiar and unfamiliar response format. If understanding appears only in the familiar form, the tutor has learned something important: the learner may need interface or format familiarisation before the unfamiliar version becomes valid evidence of the target capability.
If the examination itself uses that unfamiliar format, format fluency eventually becomes educationally relevant. The Task Purity Check does not erase real conditions; it helps the tutor separate what must be learned from what has already been learned.
7. Timing Can Turn a Knowledge Task Into a Performance Task
Adding a clock changes what a task measures. A learner who solves ten questions accurately without time pressure but completes only six under a realistic limit may have a performance-speed problem rather than a knowledge problem. Conversely, a learner who answers rapidly but inaccurately may have adopted a speed strategy that masks fragile execution.
ETS work on test speededness illustrates the broader measurement principle: timing can introduce variance that changes the interpretation of scores. In tutoring, this is why The Timed Set separates adding the clock from proving learning.
The Task Purity Check therefore asks whether the current claim is about knowledge, selection, execution or performance under time. If the tutor is diagnosing method understanding, aggressive timing may contaminate the evidence. If the tutor is preparing for real examination performance, the clock belongs to the construct.
Purity is always relative to the question being asked.
8. Background Knowledge Can Help or Distort
Contexts are not neutral. A reading passage about sailing, a Mathematics problem about foreign currency, or a Science question about an unfamiliar device may create differences that are partly about prior exposure.
Sometimes that background knowledge is legitimately part of the task. A subject expert should know the domain. Sometimes it is not. If a Primary Mathematics task intends to test proportional reasoning, obscure nautical vocabulary may add noise without adding mathematical value.
The tutor can test the possibility by preserving the mathematical structure while changing the context to something familiar. If performance improves sharply, the tutor has not proved that context was the only cause. The tutor has shown that the original result cannot safely be interpreted as pure evidence of proportional reasoning.
Later, return to varied contexts. Transfer matters. The goal is not to protect learners from unfamiliar worlds; it is to know when unfamiliarity is the thing being trained.
9. Task Purity and Accessibility Are Allies, Not Opponents
Accessibility support is often misunderstood as making a task less rigorous. A better question is whether the support changes the target capability or merely removes a barrier to showing it.
If the target is mathematical reasoning, enlarging text, providing an approved visual-access format or allowing an input method that the learner can physically operate may improve the quality of the evidence. If the target is reading fluency, having someone read the passage aloud would change the target operation and therefore change what the task can support.
This boundary is developed more fully in The Access-Support Boundary. The Task Purity Check adds the measurement perspective: a barrier that is irrelevant to the intended capability can make the evidence dirtier, not more rigorous.
Support should therefore be judged by function. What operation does it remove? What operation does it leave for the learner? What claim will the resulting performance justify?
10. Use Paired Tasks to Find the Hidden Requirement
One of the simplest diagnostic methods is to create two tasks with the same intended capability but one changed non-target demand. This is not a formal experiment. It is a disciplined comparison.
For example, Alicia receives two algebra problems with the same underlying method. One has the equation already represented. The other embeds the relationship in a paragraph. She solves the represented problem accurately but cannot form the equation from the paragraph. The tutor now has stronger reason to investigate interpretation-to-representation rather than algebraic execution.
Or Beatrice receives two inference questions. In one, the evidence lines are adjacent. In the other, the necessary evidence is distributed across the passage. If only the distributed version fails, the active problem may involve evidence integration rather than inference in the broad sense.
Paired tasks become powerful when the tutor changes one meaningful condition while preserving the underlying job. The closer the pair, the more informative the difference can become.
11. Do Not Over-Purify the Task
There is an opposite failure. The tutor strips away every difficulty until the learner performs perfectly on a task that no longer resembles the real educational demand.
A Mathematics learner may succeed when every problem is labelled by topic, the correct formula is supplied, units are pre-converted and the first step is highlighted. That result tells us very little about independent problem solving. The task has become too pure because important parts of the real construct were removed.
The correct sequence is usually diagnostic simplification followed by reintegration. Use a cleaner task to locate the capability. Repair the missing operation. Then restore the complexity that belongs to authentic performance.
This is why The Changed Question, The Mixed Set and The Full Paper exist. Learning eventually has to survive a less purified world.
12. Task Purity Changes Across Repair, Alignment and Frontier Modes
In Repair mode, cleaner tasks are often useful because the tutor is trying to locate or rebuild an early weak link. If fraction comparison is being repaired, unnecessary reading or unfamiliar notation may be reduced temporarily so the mathematical relation can be inspected.
In Alignment mode, more of the school’s authentic task structure must remain. The learner needs to participate in current work, so the tutor should know which complexity is genuinely required by the curriculum and assessment. Purity still matters diagnostically, but the route must reconnect to real school conditions quickly.
In Frontier mode, added complexity can be intentional. The tutor may deliberately combine knowledge, representation, unfamiliar context and open-ended reasoning because the target is integrated expertise. In that case, the complexity is not contamination; it is the training objective.
The same task feature can therefore be irrelevant in Repair, necessary in Alignment and desirable in Frontier. Mode clarifies what the tutor is entitled to infer.
13. Three-Student Tutorials Need Individual Purity Checks
In a three-student tutorial, one common task can be impure for different learners in different ways. Alicia may understand the language but struggle with method selection. Beatrice may solve the mathematical relation once someone clarifies the vocabulary. Ciara may understand both but lose the thread when several conditions must be held in mind.
The group therefore should not be diagnosed as a unit merely because everyone got the same question wrong. Shared output can hide different causal routes.
A tutor can preserve group coherence while making small individual discriminations. Ask Alicia to state the method. Give Beatrice the same relation in simpler language. Let Ciara annotate the conditions before solving. The goal is not to give three different curricula. It is to discover which part of the common task is performing the diagnostic work for each learner.
This is one reason small-group tutoring can be powerful when observation remains individual rather than mechanically equal.
14. A Worked Composite Case: The “Weak Algebra” Diagnosis
The following is a constructed teaching case, not a real student record.
A Secondary learner repeatedly fails algebra word problems. The initial school result suggests a broad algebra weakness. A tutor could respond with thirty more algebra questions. Instead, the tutor performs a Task Purity Check.
First, the learner solves equations that are already represented: strong accuracy. Second, the learner converts short verbal relationships into equations: moderate accuracy. Third, the learner receives longer word problems containing distractors: severe drop. Fourth, the tutor reads one long problem aloud while the learner marks givens and unknowns: performance improves. Fifth, the learner independently annotates a new long problem: performance improves again.
The evidence does not prove that algebra is perfect. It changes the diagnosis. The active weakness is more specific: extracting and representing mathematical relationships from dense language under independent conditions.
The repair plan now becomes targeted. Practise identifying relationships, moving from short to longer language, then return to mixed authentic problems. The tutor has not lowered the standard. The tutor has stopped spending the learning budget on the wrong mechanism.
15. A Worked Composite Case: The “Weak Inference” Diagnosis
Another constructed case: a learner loses marks on inference questions. The tutor notices that when evidence is located in one nearby sentence, inference is strong. When the answer requires combining a pronoun reference from one paragraph with a motive implied later, performance collapses.
A generic inference programme would be wasteful. The tutor creates paired tasks. In the first, the relevant evidence is highlighted but the inference must still be generated. The learner succeeds. In the second, evidence must be located across the passage. The learner struggles. In the third, the learner is asked only to identify which two lines belong together, without writing the inference. That step also struggles.
The cleaner evidence suggests that the active problem is evidence integration and reference tracking, not inference production alone. The tutor can now route the intervention more precisely.
This is the practical value of task purity: it turns broad labels into testable educational jobs.
16. The Task Purity Card
- Target capability: What exact operation am I trying to observe?
- Task stack: What other operations are required?
- Legitimate complexity: Which extra demands are part of the real target?
- Possible contamination: Which demands may distort the result?
- Changed condition: What one requirement can I alter without supplying the target answer?
- Comparison: Does performance change meaningfully?
- Interpretation: What claim becomes stronger, weaker or still uncertain?
- Reintegration: How will the learner return to the full authentic task?
The card should remain short. It is not a psychometric report. Its job is to prevent an avoidable category error: calling the target capability weak when the task may have failed for another reason.
17. Research Foundation and Evidence Boundaries
The assessment literature provides the strongest conceptual foundation for this volume. ETS materials on construct-irrelevant variance emphasise that score interpretation can be weakened when tasks require knowledge or operations outside the intended construct. ETS research has also examined technology-based response demands and test speededness as possible validity threats. These sources concern formal assessment, not tutoring sessions, so their terminology should be transferred cautiously.
The Australian Education Research Organisation’s Monitor Progress guide, updated 14 May 2026, emphasises checking what students understand and can apply, identifying learning gaps and adjusting instruction. Its Scaffold Practice guide similarly treats support as something selected in response to learning needs and gradually removed as proficiency develops. These guides do not prescribe a “Task Purity Check”; that is the Tutor Handbook’s practical synthesis for interpreting tutoring evidence.
Relevant assessment sources include ETS’s Validity and Fairness in Technology-Based Assessment and Validity Issues in Test Speededness. These establish why task features can alter the meaning of performance; they do not justify diagnosing an individual learner from one comparison.
18. Common Failure Modes
- Topic-label diagnosis: The page says “fractions”, therefore every error is treated as a fraction-concept error.
- Purity as simplification forever: The tutor removes authentic complexity and never restores it.
- Accessibility as cheating: A legitimate non-target barrier remains because the tutor believes harder always means more valid.
- One paired task becomes certainty: A useful contrast is treated as proof instead of evidence.
- Changing several conditions together: Language, timing, support and format all change, so the difference cannot be interpreted.
- Ignoring opportunity to learn: The tutor treats unfamiliar format as weakness even though the learner was never taught how the format works.
- Overlooking support provenance: One version appears stronger only because hidden adult or tool assistance supplied part of the target operation.
The antidote is not more testing. It is better-designed contrast and more modest claims.
19. Parent and Learner Communication
A parent may see a low mark and reasonably ask, “So is this topic weak?” A useful answer is precise without becoming evasive:
The result shows that the full task is currently unreliable. I want to separate the mathematical step from the reading and representation demands before deciding which part needs the main repair. Then we will put the parts back together under normal conditions.
The learner can also participate. Ask: “Which part of this question felt difficult before you started calculating?” Their answer is evidence, not a verdict. A learner may misidentify the cause. But the question makes task structure visible and supports metacognitive awareness.
The best outcome is not a learner who demands simplified tasks. It is a learner who can recognise whether difficulty comes from understanding the concept, interpreting the task, holding information, selecting a method, executing it or managing performance conditions—and then choose an appropriate response.
20. The Ethical Standard
Broad labels are convenient. “Weak in Science.” “Bad at word problems.” “Poor at inference.” “Careless.” They compress complex evidence into a sentence adults can remember.
That convenience carries risk. A learner can spend months practising the wrong thing because a task that required six operations was interpreted as a clean measure of one.
The Task Purity Check does not promise perfect diagnosis. It asks the tutor to earn specificity. Name the capability. Inspect the task stack. Change one plausible contaminating condition. Recheck. Preserve uncertainty. Then rebuild toward authentic complexity.
A hard task is not automatically a good diagnostic task. The tutor’s responsibility is to know what the learner had to do, what the learner actually did, and whether the evidence is clean enough to justify the label being applied.
That is the Task Purity Check.
That is Tutor Handbook Volume 0100.
Connected Reading
- The Tutor Handbook Vol No.0089 | The Diagnostic Probe
- The Tutor Handbook Vol No.0093 | The Cue Validity Check
- The Tutor Handbook Vol No.0073 | The Access-Support Boundary
- The Tutor Handbook Vol No.0076 | The Evidence Sample
- AERO | Monitor Progress
- ETS | Validity and Fairness in Technology-Based Assessment
- ETS | Validity Issues in Test Speededness