The Tutor Handbook · Volume 0076 · Series ID THB-0076
The Tutor Handbook: Complete series index.
A tutor can collect too little evidence and change the learner model too quickly. A tutor can also collect so much evidence that the lesson becomes a test centre with a teacher attached.
Both errors usually begin with a reasonable intention. The tutor wants to know whether the learner really understands, whether a repair has held, whether a weakness is recurring, whether a new routine is helping, or whether a good result was simply an unusually friendly question. The difficulty is not recognising that evidence matters. The difficulty is deciding how much evidence is enough for the decision in front of you.
Suppose Alicia solves one fresh algebra question correctly after a week of repair. Is the weak link fixed? Probably not enough evidence. Suppose she then solves thirty almost identical questions correctly in the same sitting. Do we now know much more? Perhaps less than the count suggests. The questions may all announce the same method, the correction may still be active in working memory, and the same surface structure may be doing part of the recognition work for her.
Now imagine a different sequence. Alicia solves one fresh question without a topic label. Three days later she selects the same underlying method from a mixed set. The following week she meets the idea in a changed representation. Her school paper later requires the same reasoning under ordinary classroom conditions. Four pieces of evidence can sometimes tell a tutor more than thirty repetitions because they answer different questions about the learner’s capability.
The Evidence Sample is the smallest sufficiently varied set of fresh performances that can resolve the tutor’s current uncertainty without collecting more evidence than the decision needs.
This volume is not a statistical sampling manual and it does not invent a universal number of questions. It is a professional judgement guide for tutoring. It asks how a tutor can gather evidence that is broad enough to be useful, narrow enough to be practical, and honest enough that repeated performances are not mistaken for independent proof merely because the count is large.
Quick answer
Choose the evidence sample from the decision you need to make. If the question is whether a correction can be repeated immediately, one fresh reattempt may be enough. If the question is whether learning has held, you need delay. If the question is whether the learner can recognise when the method applies, you need variation or mixing. If the question is whether a support can be faded, you need comparable work under reduced assistance. If the question is whether a route should be retired, you need evidence across more than one relevant condition.
Do not collect five versions of the same evidence and call it five independent confirmations. Do not keep testing after the decision has already become clear merely because measurement feels responsible. Do not use one dramatic performance to update a broad learner label. Use the smallest sample that separates the plausible interpretations that matter now, then return the learner to teaching, practice and independent work.
1. The evidence sample begins with a decision, not a number
A weak approach starts with a fixed quantity: five questions, ten questions, one paper, two papers, three weeks. A stronger approach starts with the uncertainty.
What exactly is the tutor trying to decide?
- Does the learner now understand the concept?
- Can the learner retrieve it after a delay?
- Can the learner identify when it applies?
- Can the learner execute accurately without the old prompt?
- Does the capability survive a changed representation?
- Was the poor result a recurring weakness or one unusual event?
- Has a maintenance item become unstable enough to reopen?
- Is the learner ready for greater challenge?
- Can the tutor responsibly reduce intervention?
These questions require different evidence. A learner can understand an explanation and still fail delayed retrieval. A learner can retrieve a procedure and still fail to select it from neighbouring methods. A learner can select the method in an untimed set and still lose the route under examination pressure. One evidence design cannot answer every version of “does the learner know it?”
This is why How Learning Diagnosis Works remains the broader diagnosis owner and The Learning Claim remains the boundary on what may be said. The present job is narrower: deciding which performances need to be sampled before that claim is updated.
2. Repetition is not the same as independence
Ten correct answers may look impressive. The evidential value depends on what changed between them.
If all ten questions are consecutive, use the same wording pattern, sit under the same chapter heading, follow the same worked example and are completed with the same support, they may primarily show that the learner can repeat one route while its cues remain highly available. That is useful. It may establish short-term execution. It does not automatically establish durable, discriminating or transferable learning.
A fresh item earns more independence when it removes some of the information that made the previous item easy to recognise. Delay removes immediate memory of the correction. Mixing removes the chapter label. New wording removes familiarity with the surface. A changed representation tests whether the learner can recover structure from a different form. Reduced prompting tests whether the learner rather than the tutor now carries the relevant decision.
The tutor therefore asks not only, “How many examples?” but, “How independent are these examples from one another as evidence about the capability I care about?”
3. Freshness has several meanings
A fresh task is not merely a question the student has never seen before. Freshness can refer to time, surface, context, method selection, representation, support or source.
Temporal freshness means the learner returns after enough time that the previous explanation is no longer sitting in immediate memory. Surface freshness means numbers, wording or presentation have changed. Contextual freshness means the same idea appears in another situation. Selection freshness means the method is not announced in advance. Representation freshness means a diagram, table, symbolic form or prose description replaces the original format. Support freshness means the learner receives less or different help. Source freshness may come from school work rather than tutor-produced material.
Not every sample needs every kind of freshness. The tutor chooses the kind that tests the claim currently under review.
4. Evidence should vary only what matters
Variation can become a second error if it changes too much at once. If the tutor wants to know whether method selection has improved, changing the method family, representation, time limit, vocabulary difficulty and support level simultaneously creates an uninterpretable result. Failure could come from any one of those changes.
Good evidence samples vary deliberately. Hold enough constant that the result can still answer the intended question.
For example, Beatrice has been learning to choose textual evidence that directly supports an inference. The tutor wants to know whether she can now discriminate between evidence that is merely related and evidence that is decisive. The next sample should contain fresh passages and questions, but it does not need unusually difficult vocabulary, a strict time limit and a completely new answer format. Those additional demands would change the question being tested.
This principle aligns with the wider formative-assessment idea that evidence should be gathered relative to clear learning goals. The OECD’s 2025 work on formative assessment and feedback describes an ongoing process of setting goals, diagnosing learning, giving feedback and adapting teaching to student thinking. A tutor’s evidence sample is one local implementation of that process, not a separate measurement system.
5. One item can be enough for a narrow decision
Tutors sometimes over-test because they fear acting on insufficient evidence. Yet some decisions are intentionally narrow.
Suppose Ciara has just corrected a Science explanation after confusing an observation with a causal mechanism. The immediate question is not whether the learning will still be available next month. The immediate question is whether she understood the correction well enough to produce a second answer without copying the original.
One well-chosen fresh reattempt may answer that local question. If Ciara succeeds, the tutor can move to the next stage: a later return. The tutor does not need ten immediate variants before arranging the delayed test.
This is a useful distinction: enough evidence for the next decision is not the same as enough evidence for a permanent learner claim.
6. Broad claims need broader samples
The stronger the claim, the more demanding the evidence sample should become.
“Denise corrected the sign error on the next question” is a narrow claim. “Denise’s sign-control weakness has been repaired” is broader. “Denise is now accurate under full-paper pressure” is broader still.
A broad claim should survive more than one relevant condition. It may require delay, different question families, ordinary school work and realistic time pressure. The goal is not maximal data. It is enough coverage of the conditions the claim itself implies.
Volume 0040, The Independence Test, is one example. Independence is not established by a single unsupported success because the claim extends beyond one moment. The evidence sample must match the breadth of the claim.
7. Samples should include the conditions under which failure matters
If a learner only needs a capability under one narrow condition, test that condition. If real performance requires the capability under several conditions, the evidence sample must eventually include them.
Alicia’s algebra method selection matters in mixed papers, not only in chapter exercises. Beatrice’s evidence selection matters in unfamiliar passages, not only in passages the tutor has discussed. Denise’s Additional Mathematics execution matters under time pressure, not only when time is unlimited. Emily’s study planning matters during busy weeks, not only during calm ones. Faith’s independent start matters when the task is slightly unfamiliar, not only when the worksheet pattern is obvious.
The tutor does not need to test every extreme. The sample should include the realistic conditions that define success for the present educational job.
8. Use short checks before large assessments when the uncertainty is small
A full paper is expensive evidence. It consumes time, fatigue and review capacity. If the tutor only needs to know whether one prerequisite is still fragile, a full paper may be an inefficient instrument.
The National Student Support Accelerator’s current example tutoring session structure includes a data touchpoint and brief formative assessment within a consistent tutoring session. Its broader data-use guidance emphasises collecting, analysing and using formative assessment data to inform future sessions. These are programme-level resources, not validation of any particular eduKate routine, but they support a sensible principle: use evidence to guide instruction without making the assessment itself the whole intervention.
Start with the smallest check that can answer the uncertainty. Escalate the evidence sample only if the first result remains ambiguous or the decision itself carries larger consequences.
9. Ambiguous evidence should trigger discrimination, not volume
A learner answers one question incorrectly. The tutor could assign twenty more questions of the same type. That produces more observations but may not resolve why the error occurred.
A better next item changes one relevant feature. Simplify the language while preserving the mathematics. Keep the wording but supply a representation. Remove the time limit. Ask for the first step only. Change the numbers while preserving structure. The goal is to separate plausible explanations.
This is the same professional logic that underlies The Differential: keep competing explanations alive until the evidence separates them. The Evidence Sample adds the question of sufficiency. Once the relevant alternatives have been separated well enough for the next decision, more testing may add little value.
10. Beware friendly evidence
Tutors naturally create tasks that match what they have just taught. This is useful for practice and dangerous for broad inference.
Friendly evidence includes questions that use the tutor’s wording, examples whose structure was rehearsed moments earlier, tasks presented under the chapter name, answer formats that mirror the worked model, or problem sets from which distracting alternatives have been removed. These conditions can help the learner acquire the route. They are not automatically appropriate conditions for proving that the learner can select or transfer the route independently.
eduKateSG’s Study Selection Bias explores a related general problem: friendly practice can create false readiness. The Tutor Handbook question is operational. When the tutor is ready to update the learner model, include at least some evidence that was not curated to make the newly taught method obvious.
11. Beware hostile evidence too
“Fresh” should not mean deliberately obscure. A tutor can make the sample so unfamiliar, language-heavy or time-pressured that it begins measuring additional demands rather than the capability of interest.
Suppose Faith is learning to start multi-step tasks independently. Giving her a problem containing unfamiliar technical vocabulary may produce hesitation, but the hesitation does not cleanly show that independent task initiation failed. The language itself became a new barrier.
An honest evidence sample includes appropriate challenge rather than surprise for its own sake.
12. Sample support conditions explicitly
Evidence becomes difficult to interpret when the amount of help is invisible.
A correct answer after a method cue is different evidence from a correct answer after a task-restatement cue, and both differ from an unsupported answer. Legitimate access support should remain in place where it is not the skill being assessed. Volume 0073, The Access-Support Boundary, owns that distinction.
The Evidence Sample asks for declared support conditions so performances can be compared honestly. If the tutor wants to know whether a scaffold can fade, sample the same intellectual job with the scaffold present and later with a carefully reduced version. Do not remove unrelated accessibility support merely to make the test look independent.
13. School work can provide valuable independent evidence
Tutor-made tasks have one advantage: they can be designed precisely around the current question. School work has another: it is produced outside the tutor’s immediate instructional environment.
A marked school paper, classroom task or teacher comment can therefore become useful evidence that a capability travelled. It should not be treated as automatically superior. The paper may test different content, permit different support, use a different marking scheme or occur under unusual conditions. But when the conditions are understood, school evidence can broaden the sample beyond tutor-curated work.
The National Student Support Accelerator’s guidance on teacher–tutor communication explicitly treats formative and summative classroom data as useful inputs for tutoring. For an independent tutor, access will vary and privacy boundaries matter. Use only information legitimately available to the learner or family, and never assume the school task measured exactly the same construct as the tuition task.
14. A three-part evidence sample often works better than a long same-session set
For many tutoring decisions, a useful starting design is not “do ten questions” but three different receipts:
- Immediate receipt: Can the learner make the corrected or newly taught move on a fresh attempt now?
- Delayed receipt: Can the learner recover the move later without the explanation still active?
- Changed-condition receipt: Can the learner identify or use the capability when the surface, context, support or neighbouring alternatives change?
This is not a universal protocol and should not be applied mechanically. Some tasks require more observation; others require less. Its value is conceptual: it reminds the tutor that evidence quality can come from independence across conditions rather than sheer quantity in one sitting.
15. Worked case: Alicia’s algebra method selection
This is a fictional teaching case used to demonstrate the decision process.
Alicia has been choosing the wrong algebraic method when different question families are mixed. The tutor runs a targeted repair and Alicia then solves six similar questions correctly.
The wrong conclusion is: “Six out of six; method selection is fixed.” The six questions were blocked by type, so they mainly showed execution after the method had already been identified by the set.
The tutor needs a different evidence sample. On the next session, Alicia receives a short mixed set containing three familiar method families. She must name the relationship she sees before working. She selects four of five correctly. Three days later, she receives two changed-wording items with no topic label and selects both correctly. Her next school assignment includes one comparable mixed item, which she also handles correctly without a tutor prompt.
Notice what changed. The tutor did not collect twenty more algebra questions. The tutor collected evidence under conditions relevant to the claim: mixed selection, delay, changed surface and outside-tuition performance. That is enough to justify moving method selection from active repair toward maintenance, while still keeping a small watch rather than declaring permanent mastery.
16. Worked case: Beatrice’s comprehension evidence
This is a fictional teaching case.
Beatrice has learned to distinguish a direct textual clue from a merely related sentence. During the lesson she answers five tutor-written questions correctly. The passages are short and the relevant clues sit close to the questions.
The tutor wants to know whether evidence selection is stable enough to reduce explicit prompting. Instead of another large same-format set, the tutor chooses two longer fresh passages on different days. In the first, Beatrice must choose between two plausible supporting lines. In the second, the relevant clue is separated from the inference by several sentences. She explains why one candidate line is stronger than another.
She succeeds on the first and becomes over-general on the second. The evidence sample has done its job: it has revealed a boundary. The tutor does not need a verdict of “mastered” or “not mastered.” The learner model becomes more precise: direct evidence selection is stable in local contexts; distal evidence integration still needs work.
17. Worked case: Denise and the danger of a full-paper overreaction
This is a fictional teaching case.
Denise loses several marks on one Additional Mathematics paper through sign errors. A tutor could immediately reopen full sign-control repair. But the paper occurred after a particularly dense school week and contained several unfamiliar question structures.
The tutor chooses a smaller evidence sample first: a short representative section completed under ordinary timing, then another section several days later. If the old sign pattern recurs across both, reopening becomes more justified. If the errors disappear while other difficulty remains, the tutor avoids rebuilding an old intervention from one noisy event.
This is evidence efficiency. The tutor spends enough measurement to protect the learner from a false conclusion, but not so much that a full paper becomes the automatic response to every uncertain signal.
18. Stop when the decision is resolved
Good measurement has a stopping rule.
If the tutor’s uncertainty was whether the learner could retrieve a repaired concept after three days, and the learner demonstrates it clearly under the planned condition, there may be no benefit in adding eight more immediate tests. Move to the next learning job and schedule another appropriate return later if the claim requires durability.
Testing can become attractive because it produces visible information. Teaching and practice often produce slower, messier information. The tutor should remember that the purpose of evidence is to improve the next learning decision, not to maximise the amount of data collected.
19. Over-testing has educational costs
Every assessment opportunity displaces something else. It uses lesson time, attention, emotional energy and practice capacity. Frequent checking can also change the relationship with the subject if the learner feels that every action is being judged rather than used for learning.
This does not mean frequent formative checking is bad. AERO’s Monitor Progress guide describes regular checking for understanding as a way to determine what students know and can apply and to provide additional instruction, guidance or feedback. The important word is responsive. Checks should change teaching when they reveal something useful. When the result would not change the next action, the tutor should ask whether another check is earning its time.
20. Under-testing has costs too
The opposite error is to accept the first encouraging result because everyone wants to move on.
A same-session correction can feel like a breakthrough. A fluent explanation can feel like mastery. A single high score can feel like proof. If the decision is consequential—ending repair, removing support, moving into substantially harder work or releasing the learner from routine intervention—the evidence sample should usually include conditions that test the claim’s weak points.
Volume 0074, The Challenge Readiness Gate, asks whether harder work will extend learning or hide unstable foundations. The Evidence Sample is one way of earning that decision rather than guessing it.
21. Group tutoring needs individual evidence samples
In a three-student tutorial, one shared task can produce three different kinds of evidence.
Alicia may need evidence about method selection. Beatrice may need evidence about explanation quality. Ciara may need evidence about scaffold independence. The tutor should not assume that because all three completed the same worksheet, the worksheet answered the same professional question for each learner.
This is one reason small-group tuition can be powerful when the tutor can see individual thinking, and weak when the group becomes one averaged learner. The EEF’s current small-group tuition summary notes that targeting support to specific pupil needs matters. It does not validate a particular three-student protocol, but it reinforces the importance of preserving individual need inside group delivery.
22. Parents should ask what the evidence actually changed
A parent may see several worksheets and ask whether the child has improved. A better conversation begins with the decision those worksheets were meant to inform.
- What were we uncertain about?
- What condition did the fresh work change?
- Was the learner working independently?
- Did the result repeat after a delay?
- Did it appear outside the exact practice format?
- What did the tutor change because of the evidence?
The last question is especially useful. If the evidence changed nothing about teaching, practice, support or monitoring, the family should ask whether the evidence was collected because it was needed or because assessment had become a habit.
23. The learner should increasingly understand the sample
Evidence should not remain an adult surveillance system. As learners mature, they can understand why different checks exist.
“You got these right immediately. I want to see whether the idea is still available on Friday.”
“You can solve these when the chapter tells you the method. Now I want to see whether you can choose the method when several are possible.”
“We are not testing you again because the first result was bad. We are checking whether that result is a recurring pattern before we reopen a whole repair route.”
These explanations turn assessment into metacognitive information. The learner begins to see why one success is not the same as durable learning and why one failure is not a permanent identity.
24. A practical Evidence Sample card
- Decision: What exact decision will this evidence inform?
- Claim breadth: How broad is the conclusion we hope to make?
- Immediate evidence: What can be checked now?
- Delay: Does the claim require later retrieval?
- Variation: Which surface or context should change?
- Selection: Must the learner decide when the method applies?
- Support: What help is permitted and recorded?
- Outside evidence: Would school or independent work add useful information?
- Alternative explanations: What else could produce success or failure?
- Stopping rule: What result would be enough to make the next decision?
- Next return: If the claim requires durability, when should the capability be sampled again?
25. What this volume does not authorise
The Evidence Sample does not authorise tutors to invent psychometric scores, diagnose clinical conditions, treat one tutoring sample as a standardised assessment, or claim statistical certainty from small informal sets. It does not say that more testing is always better or that a fixed three-check sequence proves mastery. It does not remove legitimate access support. It does not override school assessment rules or formal examination standards.
It is a practical discipline for everyday tutoring: choose evidence that is fit for the decision, vary the conditions that matter, keep the claim proportional to the sample, and stop measuring when the evidence has done its job.
Research foundation and limits
The direct evidence base supports the broader ingredients rather than this exact Tutor Handbook framework. AERO’s Monitor Progress guidance emphasises frequent checking for understanding, identifying gaps and adjusting teaching. The OECD’s formative assessment and feedback chapter describes an ongoing process of eliciting student thinking and adapting instruction. The National Student Support Accelerator’s data-use guidance emphasises formative assessment, structured review and the use of data to personalise tutoring. EEF’s small-group tuition review highlights targeting support to specific needs.
Those sources do not establish a universal sample size for private tutoring, nor do they validate the examples above as a tested programme. The practical contribution of this volume is to connect established formative-assessment principles to one recurrent tutor decision: how to collect enough fresh evidence without confusing quantity with quality or assessment with learning.
Final principle
A tutor does not need the largest possible sample.
The tutor needs the smallest sample that is honest about the decision.
If the decision is narrow, sample narrowly. If the claim is broad, sample the conditions the claim includes. When evidence is ambiguous, change the question rather than merely increasing the count. When the learner has proved the point, stop measuring and return to learning.
The Evidence Sample is enough fresh work to make the next judgement better—no less, and no more than the learner needs to carry.