The Tutor Handbook · Volume 0095 · Series ID THB-0095
The Tutor Handbook: Complete series index.
A learner has just watched the tutor solve a percentage problem.
The tutor removes the worked example, changes 20% to 15%, and asks for another attempt.
The learner succeeds.
What does that success tell us?
Something useful. But not everything.
The route is still warm. The language, representation, method, correction and tutor emphasis are all close enough to the new question that the learner may be carrying part of the previous solution forward. That is not cheating. It is normal learning. The mistake begins when a tutor treats this immediate supported-neighbour success as though it were the same evidence as a later fresh performance.
The Evidence Freshness Window is the tutor’s judgement about whether a performance is sufficiently separated from recent teaching, hinting, rehearsal, answer exposure or peer support to serve the particular learning claim being made.
This is not a stopwatch rule. There is no universal number of minutes after which an answer becomes “independent”. Freshness depends on what the learner was shown, how similar the next task is, what cues remain available, what skill is being claimed, and how much the learner must reconstruct rather than recognise.
Quick answer
Immediate success is valuable instructional evidence. It can show that the learner followed an explanation, used feedback, corrected an error or completed a newly modelled procedure. But stronger claims—retention, independent method selection, transfer, self-regulation or durable understanding—need evidence that is fresher than the teaching event that produced the success.
A tutor should therefore ask two separate questions:
- Did the learner improve on the next attempt?
- Was that attempt fresh enough to support the claim I want to make?
The first question protects responsive teaching. The second protects honest interpretation.
1. What this volume owns
This volume owns the evidence boundary between a recently taught or assisted response and a sufficiently fresh response for a stronger learning claim.
It does not own the general mechanism of spacing. AERO’s Vary Practice guide and eduKate learning owners already cover spaced and varied practice.
It does not own evidence returning from outside the lesson. Volume 0017 | The Return owns that broader question.
It does not own the second attempt itself. Volume 0018 | The Reattempt owns how a tutor reads a second attempt.
It does not own changed-surface transfer. Volume 0019 | The Changed Question owns that test.
It does not own support provenance. Volume 0078 | The Support Provenance Check owns who or what helped produce the work.
This article connects those owners around a narrower practical decision: is this response still too close to the teaching event to carry the claim I am about to place on it?
2. Performance can improve before learning has stabilised
The distinction between performance during acquisition and more durable learning is well established in learning research. Soderstrom and Bjork’s review, Learning Versus Performance, explains why conditions that produce strong immediate performance do not always produce the strongest long-term learning, and why performance observed during practice can misrepresent what will remain later.
That does not mean immediate performance is unimportant. A tutor needs it constantly. It tells us whether an explanation landed, whether a prompt was understood and whether a correction can be used now.
The point is narrower: the evidential job of an immediate response differs from the evidential job of a delayed, changed or unsupported response.
3. Warm evidence is still evidence
It is useful to call the response immediately after teaching warm evidence. This is a descriptive teaching term, not a validated scientific scale.
Warm evidence can answer questions such as:
- Can the learner follow the newly explained route?
- Can the learner use the feedback on the next attempt?
- Did the tutor’s clarification resolve the immediate misunderstanding?
- Can the learner reproduce the new procedure with the structure still recently active?
- Did a scaffold make the target operation possible?
These are important questions. A tutor who refuses to look at warm evidence would teach blindly.
The error is not using warm evidence. The error is using it to answer a colder question.
4. Fresh evidence asks for reconstruction
Fresh evidence becomes more informative when the learner has to reconstruct enough of the route that the recent teaching event can no longer do most of the organising.
Freshness can increase through several kinds of separation:
- time: the learner returns later rather than immediately;
- surface: the wording, numbers, representation or context changes;
- selection: the topic label or method cue disappears;
- support: a scaffold, model answer or peer explanation is no longer present;
- sequence: other material intervenes so the learner must choose the route again;
- responsibility: the learner, rather than the tutor, initiates the next move.
No single dimension guarantees freshness. A task can be delayed but nearly identical. A task can look different while still containing a strong cue. A response can be unsupported but follow an answer the learner memorised seconds earlier.
5. Freshness belongs to the claim
The same performance can be fresh enough for one claim and too warm for another.
A learner solves a structurally similar equation two minutes after guided practice.
That may be fresh enough to claim:
The learner can now execute the procedure with the worked example removed.
It is not necessarily fresh enough to claim:
The learner has retained the procedure and can select it independently in mixed work.
Evidence freshness is therefore not an intrinsic property of the worksheet. It is a relationship between the observed performance and the conclusion the tutor wants to draw.
6. The strongest contamination is often cue persistence
A learner may no longer see the answer but still retain the route as an active cue.
The tutor says, “Remember: identify the base first.”
The learner succeeds on the next percentage question.
The paper contains no written help, but the relevant decision has just been named.
For a claim about calculation, that may be acceptable. For a claim about independently identifying the percentage base, the evidence is still warm.
The tutor should ask: Which part of the target operation did the recent cue already perform?
7. A changed number is not necessarily a changed problem
Changing 80 to 120 does not create meaningful freshness if the learner can simply replay the same sequence.
For example, after modelling “What is 25% of 80?”, the tutor asks “What is 25% of 120?” The second item is useful practice. But it barely tests method selection.
A stronger freshness step might ask: “18 is 25% of what number?” Now the same percentage relationship appears with a different unknown. The learner must recognise the structure rather than merely substitute a new number into the same visible route.
8. A new surface can still contain an old answer
English tutoring creates a similar problem.
The tutor models how to infer hesitation from a character’s behaviour, using the evidence “she paused at the doorway and looked back”.
The next passage says a character “stopped outside the room and checked his phone twice before entering”.
The surface is new, but the semantic pattern is extremely close. Success is encouraging. It is not yet broad evidence that the learner can infer attitude or emotion from unfamiliar textual evidence.
Freshness rises when the learner has to identify a different evidence pattern, explain why it supports the inference, and reject an attractive alternative.
9. Model-answer language can remain active after the model disappears
Science tutoring often uses strong model explanations. The tutor shows a causal chain, hides it and asks the learner to answer a nearly identical question.
The learner produces the same phrases.
This may show successful short-term reproduction. It does not yet show that the learner can reconstruct the mechanism.
A fresher check changes the condition: alter one variable, change the representation, ask for a prediction, or ask the learner to explain which part of the causal chain would fail if one condition changed.
The goal is not to make the task artificially difficult. It is to require the learner to regenerate the relationship rather than recall the sentence.
10. The tutor’s own words can linger
After a good explanation, the learner may sound unusually fluent because the tutor’s language is still available in working memory.
This is particularly easy to misread in one-to-one and small-group tuition, where the tutor hears the learner use an elegant phrase almost immediately after saying it.
Do not dismiss the response. Instead, label it accurately:
Accurate explanation immediately after modelling; independent reconstruction not yet checked.
That sentence preserves both the success and the uncertainty.
11. Peer discussion changes freshness
In a three-student tutorial, information travels between learners. One student names the method; another explains the evidence; a third repairs the misconception.
Volume 0094’s Peer Answer Leakage Boundary protects the meaning of independent first responses before discussion. After discussion, the final answer may be excellent—but it is no longer fresh evidence of what each learner could do before the group supplied information.
Later individual return restores a different evidence condition.
12. AI can make an answer look fresher than it is
A learner completes homework at home, then brings an excellent explanation to tuition. The learner says they “only checked it with AI”.
Checking can range from spelling correction to receiving a complete causal explanation.
The tutor should not argue about the word “checked”. Use Volume 0078’s support-provenance question: what operation did the tool perform?
Then create a fresh task that requires the target reasoning without the same assistance. The issue is not technology morality. It is evidential clarity.
13. Freshness does not require withholding feedback
A bad response to this problem is to delay teaching so that every attempt remains diagnostically pure.
That reverses the purpose of tuition.
If the learner needs an explanation, teach.
If the learner needs feedback, give it.
If the learner needs a scaffold, provide the minimum support that allows the target operation to develop.
Then change the label on the evidence. Teaching can continue now; stronger independent evidence can be collected later.
14. The instructional loop and the evidence loop are related but not identical
The instructional loop may be fast:
Attempt → feedback → correction → immediate new attempt.
The evidence loop may need more separation:
Teach → practise → remove or reduce support → change condition → return later → inspect what survives.
A tutor can run both loops without conflict once their purposes are separated.
15. A practical freshness ladder
A useful way to organise evidence is to increase separation gradually.
- Same-item correction: Can the learner repair the exact error?
- Near-neighbour attempt: Can the learner use the new route on a similar item?
- Changed-surface attempt: Can the learner recognise the same underlying relationship when the surface changes?
- Mixed selection: Can the learner choose the route when no label announces it?
- Delayed return: Can the learner reconstruct the route after the immediate teaching context has faded?
- Independent transfer: Can the learner use the capability in a new task, school paper or real performance condition?
This is not a compulsory six-stage protocol. Some skills do not need every rung. The ladder is a way to ask whether the next piece of evidence answers a stronger question than the previous one.
16. Composite case: Alicia and algebra
The following case is illustrative, not a claim about a real student.
Alicia incorrectly expands 3(x + 4) as 3x + 4. The tutor uses an area representation and then explains the distributive relationship.
Alicia immediately corrects the original item.
Evidence job: she can follow the correction.
She then expands 5(y + 2) correctly.
Evidence job: she can reproduce the structure on a near neighbour.
Later, after working on equations and graphs, she sees -2(3a – 1) in mixed work and expands it correctly without a prompt.
Evidence job: the distributive relationship is available under a fresher selection condition.
The tutor does not need to call the first success “not learning”. Each success simply supports a different claim.
17. Composite case: Beatrice and comprehension evidence
Beatrice chooses a broad paragraph as evidence for a specific inference. The tutor asks, “Which exact phrase most directly supports the inference?”
Beatrice selects the right phrase immediately.
That is evidence that the cue worked.
If the tutor then asks another question using the same paragraph and says, “Find the exact phrase again,” the critical evidence-selection decision is still strongly cued.
A fresher check uses a new passage and does not remind Beatrice to hunt for the narrowest line. If she independently chooses precise evidence, the tutor can update the learner model more confidently.
18. Composite case: Ciara and Science causal explanation
Ciara gives a correct conclusion but omits the mechanism. The tutor models a cause-and-effect chain, then asks Ciara to repeat it.
Her repetition matters: she has heard and represented the mechanism accurately.
Next, the tutor changes one condition and asks what would happen. Ciara must now use the mechanism to generate a prediction.
A later delayed question presents a different context with the same underlying relationship. If Ciara reconstructs the causal chain again, the evidence is fresher and the learning claim can expand.
19. Composite case: Emily and studying
Emily struggles to prioritise a busy study week. The tutor models a simple rule: fixed obligations first, then high-value weak links, then maintenance, with a buffer left open.
Emily builds a good plan while the tutor is beside her.
That is not yet evidence of self-directed planning.
The following week, Emily brings a plan she created before tuition. Now the tutor can inspect whether the same decision principles survived without live adult selection.
Freshness here is partly temporal and partly a transfer of responsibility.
20. Freshness and access support must be separated
A learner may legitimately use enlarged text, text-to-speech, a screen reader, approved notation support or another access accommodation. Removing those merely to create “fresh” evidence can make the task less fair rather than more independent.
Ask whether the support performs the target intellectual operation.
If the target is algebraic method selection, large print does not select the method.
If the target is reading decoding, text-to-speech may materially change the target.
Freshness is therefore always construct-specific. Preserve legitimate access support while removing only the answer-giving or target-performing assistance relevant to the claim.
21. Freshness and confidence are different
A learner can answer confidently because the solution was just discussed.
A learner can answer hesitantly on a genuinely fresh task and still be correct.
Do not use confidence as a substitute for freshness. Volume 0093’s cue-validity principle applies: ask whether the cue actually bears on the target claim.
22. Freshness and difficulty are different
A harder question is not automatically fresher evidence.
The tutor can make a task difficult by adding irrelevant complexity while preserving the same cue.
For example, a percentage question can contain more numbers and more words yet still announce the exact operation. Difficulty rises; method-selection freshness does not.
Change the dimension that matters to the claim.
23. Freshness and randomness are different
Randomly selecting questions from a worksheet does not guarantee that the tasks differ meaningfully from the taught examples.
Question selection should preserve the target capability while changing the support condition in a purposeful way.
Randomness can be useful for reducing predictability, but it is not a substitute for good task design.
24. Freshness and delay are different
Waiting a week and repeating the exact distinctive question can still produce recognition.
Conversely, a changed problem later the same lesson may reveal meaningful transfer.
Delay is one source of freshness, not the whole definition.
AERO’s current Vary Practice guide supports varied and spaced opportunities as part of building adaptable knowledge and using practice to monitor progress. The guide does not prescribe one universal delay for every learning claim.
25. Freshness should increase with the strength of the claim
A modest claim needs modest separation.
“The learner used today’s feedback correctly on the next attempt” can be supported immediately.
“The learner has mastered the skill” is stronger. It should usually survive a greater change in time, task, support and selection condition.
“The learner can regulate this process independently” is stronger again. The learner should initiate, monitor and repair under conditions where the tutor no longer carries those functions.
The evidence burden rises with the claim.
26. Use the smallest freshness change that resolves the uncertainty
Do not turn every check into a full paper.
If the uncertainty is method selection, remove the topic label and mix two neighbouring problem types.
If the uncertainty is retention, return after a useful delay.
If the uncertainty is scaffold dependence, reduce or remove the relevant scaffold.
If the uncertainty is transfer, change the surface while preserving the underlying structure.
If the uncertainty is self-regulation, let the learner decide the next step before the tutor intervenes.
Freshness should be designed around the unresolved question.
27. The freshness tag
A tutor can keep a small note beside evidence:
- W — Warm: immediately after teaching, correction or strong cue.
- N — Near-fresh: changed item with recent route still likely active.
- F — Fresh: enough separation for the specific current claim.
- R — Return: later evidence from a delayed or external condition.
These labels are not validated scores and should not become a new bureaucracy. They are prompts for honest interpretation. The important sentence is always the plain-language description of the condition.
28. Parent communication: “He got it in tuition”
A parent sees a correct tuition worksheet and understandably asks whether the problem is fixed.
A useful tutor response is:
He can use the corrected method immediately now, which is a good first step. I have not yet treated that as durable independent learning because the explanation is still recent. I will check it again under a changed and later condition before moving the topic fully into maintenance.
This protects the good news without overstating it.
29. Learner communication: the return is not a trap
Learners can experience delayed checks as though the tutor is trying to catch them forgetting.
Explain the purpose:
I already know you could do it after we worked on it together. I want to see what stayed with you when the example is no longer doing part of the work. That tells us what to keep practising and what we can safely move on from.
The return becomes information rather than judgement.
30. Tutor coaching: audit claims, not only questions
When reviewing another tutor’s lesson, it is easy to ask whether enough checks occurred.
A stronger coaching question is:
What claim did you place on each check, and was the response fresh enough for that claim?
A tutor may have excellent questioning but still update the learner model too aggressively from warm evidence. Conversely, a tutor may be appropriately cautious and fail to recognise when later evidence has become strong enough to move the route forward.
31. Repair mode
In Repair mode, warm evidence is especially useful because the tutor needs fast feedback about whether the repair is operating.
But exit from Repair should normally require fresher evidence than entry into the repair. The learner should show the repaired capability after support has reduced and the original teaching event is less available.
32. Alignment mode
In Alignment mode, school tasks can provide naturally fresh evidence because they arrive from another source and often use different wording, timing and context.
Do not assume school work is automatically independent. Check for homework support, answer keys, peer discussion and tool use when those conditions matter to the claim.
33. Frontier mode
In Frontier mode, freshness protects against mistaking rapid pattern pickup for robust transferable capability.
A strong learner may reproduce a sophisticated method immediately after one demonstration. The next useful question is whether the learner can identify when the method applies, explain why it works and use it under a changed surface without the original cue.
34. Three-student tutorials
Small groups can generate rapid warm evidence because learners hear several explanations and see several attempts. That is an instructional advantage.
It also means evidence can warm quickly. A learner’s final response may contain tutor cues and peer ideas even when the learner writes independently.
Use short protected first responses when individual starting evidence matters, then let discussion flow. Later individual returns can establish fresher evidence without suppressing the benefits of group learning.
35. Common failure: “correct twice means mastered”
Two immediate correct answers may still be two performances under nearly identical conditions.
Repetition raises confidence only when the repeated evidence adds something new. Change time, surface, support or selection condition when the claim requires it.
36. Common failure: waiting without changing the task
The tutor delays a week but repeats the same distinctive worksheet item.
The learner recognises it.
Retention may indeed have occurred, but the result says less about flexible use than a fresh related problem would.
Match the return to the claim.
37. Common failure: making the check so different that it tests something else
Freshness is not maximal difference.
If the tutor changes vocabulary, representation, prerequisite load and time pressure simultaneously, failure becomes hard to interpret.
A fresh task should preserve the target operation while changing enough of the surrounding condition to answer the current uncertainty.
38. Common failure: “no help” becomes the definition of independence
Independence is target-specific. A learner may legitimately use a calculator, formula sheet, dictionary, access accommodation or approved reference while still performing the intended reasoning independently.
Remove help only when that help performs the target operation being evaluated.
39. Common failure: fresh evidence becomes continuous testing
A tutor who chases fresh evidence after every teaching move can starve the learner of actual instruction.
Teaching deserves uninterrupted stretches. Practice deserves repetition. Explanation deserves time.
Collect the smallest evidence needed to decide the next route. Volume 0076’s Evidence Sample remains relevant: more evidence is not automatically better evidence.
40. The Evidence Freshness card
- Claim: What exactly am I trying to infer?
- Recent teaching: What explanation, model, correction or cue was just provided?
- Support: What help remains available?
- Similarity: How close is this task to the taught example?
- Selection: Does the task announce the method or relationship?
- Time: Has enough context faded for the claim being made?
- Intervening work: Has the learner had to choose among other routes?
- Peer/tool exposure: What relevant information crossed into the learner’s route?
- Freshness decision: Is this warm evidence, a near-fresh check or fresh enough for this claim?
- Next receipt: What later or changed condition would strengthen the conclusion?
41. Research foundation: monitor progress without confusing performance with durable learning
AERO’s Monitor Progress guidance emphasises checking what students understand and can apply, using that evidence to identify gaps and adjust instruction, guidance or feedback.
AERO’s Scaffold Practice guidance emphasises planned and responsive support and gradually removing scaffolds as proficiency develops.
Together with the broader learning-versus-performance literature, these sources support the practical discipline behind this volume: observe current performance, respond to it, and avoid treating success under one support-rich condition as proof of a stronger capability claim than the evidence can carry.
42. Research boundary
There is no validated universal “freshness window” measured in minutes, questions or days. This article uses the term as a tutoring decision concept, not a psychometric scale.
The research base supports distinctions among immediate performance, delayed retention, varied practice, scaffolding and independent application. It does not validate one exact eduKate sequence for all ages, subjects or three-student tutorials.
Freshness should therefore be treated as a structured judgement: identify the claim, identify the support and similarity surrounding the performance, and choose a later or changed condition when the current evidence cannot yet carry the conclusion.
43. The ethical standard
Learners deserve credit for immediate success.
They also deserve protection from adults building a false model of capability from success that the teaching context partly supplied.
Underclaiming can hold a learner back. Overclaiming can remove support too early, move the route too quickly or create avoidable failure later.
The ethical middle is accurate language:
Say what the learner could do under the condition that actually existed. Then earn stronger claims by changing the condition carefully enough to see what the learner can carry forward.
Evidence and connected reading
- The Tutor Handbook | Complete Series Index
- The Tutor Handbook Vol No.0017 | The Return
- The Tutor Handbook Vol No.0018 | The Reattempt
- The Tutor Handbook Vol No.0019 | The Changed Question
- The Tutor Handbook Vol No.0076 | The Evidence Sample
- The Tutor Handbook Vol No.0078 | The Support Provenance Check
- AERO | Monitor Progress
- AERO | Scaffold Practice
- AERO | Vary Practice
- Soderstrom & Bjork | Learning Versus Performance
Final compression
Teach when teaching is needed.
Use the next attempt.
Call it what it is.
Immediate success can show uptake.
Near-neighbour success can show use.
Changed conditions can show selection.
Delayed return can show availability.
Fresh independent performance can support stronger learning claims.
Do not wait for a magical number of minutes.
Ask what the learner was recently shown.
Ask what the task still gives away.
Ask what support remains.
Ask what claim you want to make.
Then create only enough separation to make that claim honest.
The Evidence Freshness Window protects both sides of good tutoring: it lets the tutor teach responsively in the moment, while requiring stronger evidence before immediate success is promoted into a claim about durable, independent learning.
That is the Evidence Freshness Window.
That is Tutor Handbook Volume 0095.