Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

The Tutor Handbook Vol No.0192 | The Progress-Review Window Gate — How a Tutor Chooses How Much Recent Evidence to Read Together Before Changing the Learning Route Without Letting One Bad Lesson or an Old Average Dominate

The Tutor Handbook · Volume 0192 · Series ID THB-0192

The Tutor Handbook: Complete Series Index

A learner can look weak over the year and strong now

A tutor opens the learner record.

The term average is 61 per cent. The last three checks are 76, 79 and 81. One school test from six weeks ago is still pulling the average down. The learner has also changed method, reduced prompting and completed two fresh questions successfully.

What is the current learner state?

A different record creates the opposite problem. A learner has months of strong work, then two poor sessions. The latest paper is unusually weak. Should the route change immediately, or is the tutor looking at a short disturbance inside a stable pattern?

Tutors often say they "look at the data". The difficult part is deciding which data belong in the current decision.

A very short window can overreact to noise. A very long window can average away real change. A window that mixes supported and unsupported work, different task families or different access conditions can create a smooth number that represents no coherent capability at all.

The Progress-Review Window Gate asks how much recent evidence a tutor should read together before changing a learning route, how older evidence should remain useful without dominating the present, and what to do when the learner's conditions have changed enough that the historical series is no longer directly comparable.

This is not a rule that three data points are always enough, or that four weeks is the correct period. The right window depends on the decision, the target capability, the rate at which change could plausibly occur, the noise in the evidence and the cost of acting too early or too late.

Quick answer

Choose the evidence window by the decision horizon, not by convenience.

If the decision is whether to change tomorrow's explanation, recent task-level evidence may be sufficient. If the decision is whether a repaired prerequisite is stable enough to leave active Repair, use repeated fresh evidence across enough delay and variation to support that claim. If the decision is whether a term-long route has improved examination performance, include broader evidence across realistic conditions.

Start with the newest comparable evidence. Add older evidence until it materially improves interpretation. Stop extending the window when older samples describe a meaningfully different learner state, support condition, task definition or curriculum phase.

Do not allow one bad lesson to trigger a major route change unless its consequence is serious enough that waiting is unsafe. Do not allow an old average to suppress a clear new pattern simply because the average contains more observations.

When the evidence conflicts, preserve the conflict. Ask whether it reflects task variation, support differences, a recent intervention, a temporary condition, genuine instability or an actual regime change in the learner's capability.

A review window should answer the question: what evidence is still close enough in time and condition to inform this decision now?

This is not another evidence-sampling article

The Tutor Handbook already has an Evidence Sample that asks how much fresh work is enough to update the learner model without turning tuition into continuous assessment.

The Construct-Coverage Check asks whether a progress check samples enough of the target capability before a broad mastery claim is made.

The Correlated-Evidence Trap asks whether several apparent confirmations are really independent enough to count as multiple pieces of evidence.

The Decision-Threshold Gate and Precommitted Route-Change Threshold own the evidence condition that triggers, holds or reverses a decision.

The Progress-Review Window Gate owns a different dimension: time.

Which observations should be treated as describing one current state? When does evidence become stale? When does a recent shift deserve more weight than a long history? When should a tutor deliberately widen the window because the newest result is too noisy to stand alone?

A learner model has a temporal resolution. This article owns that resolution.

Current tutoring guidance makes structured data review a central job

The Stanford National Student Support Accelerator's current Utilize Data section places ongoing formative assessment, structured data reviews and longer-term progress tracking inside high-quality tutoring. It recommends that tutors use student data to personalise instruction and adjust in real time, and that programmes provide time and support to analyse evidence rather than merely collect it.

AERO's current Monitor Progress guidance, updated 14 May 2026, similarly treats progress monitoring as a cycle: gather evidence frequently enough to inform teaching, interpret it, respond and continue checking.

These are programme and practice guides, not prescriptions for one universal temporal window.

The gap appears because "review the data" is not operational until the tutor decides what period and conditions count as current evidence. A weekly score, a term average, the last three attempts and the last comparable unseen task can all be legitimate summaries. They answer different questions.

The Tutor Handbook needs a gate for choosing between them.

The newest result deserves attention, not automatic authority

Recency matters because tutoring aims to change the learner.

If teaching has worked, the newest evidence should often differ from the oldest. A system that permanently gives equal weight to every historical result can become slow to recognise improvement.

But recency is not proof of a new state.

The latest result can be unusually high because the task matched practice closely. It can be unusually low because the learner misunderstood one instruction, was interrupted, encountered a harder paper or lost access to a legitimate support. It can be affected by ordinary measurement variation.

A tutor should therefore treat the newest result as a candidate update.

The first question is not, "Is this the most recent score?" It is, "What changed in the evidence-generating condition, and does this result belong to the same underlying capability we are trying to track?"

Recency raises priority for inspection. Comparability and repetition determine how much the route should move.

Old evidence can be accurate and still be stale

Stale evidence is not necessarily bad evidence.

A marked paper from three months ago may accurately show what the learner could do then. The problem is using it as though it still describes the current state after substantial teaching, practice, support change or curriculum progression.

Staleness therefore depends on change, not merely age.

A spelling pattern that has been stable for years may still be informative after two months. A fragile algebra prerequisite can change substantially after two focused sessions. A learner's speed under timed conditions can improve after repeated realistic practice. An access arrangement may change the meaning of older timed results immediately.

Ask what has happened since the old sample.

Has there been a targeted intervention? A major change in support? A curriculum transition? A long absence? A tutor change? A new school method? A substantial period of independent practice? If yes, older evidence may remain historical context but lose authority over the current route.

The Fresh Look already protects against treating old learning history as a verdict. The review-window gate makes the temporal consequence explicit.

A progress window should be homogeneous enough to mean something

A tutor sees five recent percentages: 68, 72, 75, 77, 79.

The line looks beautifully upward.

But the first two scores came from unsupported unseen questions. The next three came from worksheet sets with method labels and immediate hints.

The trend is not a clean learning trend.

A useful review window needs enough comparability that the samples can be interpreted together. They do not have to be identical. In fact, changed conditions are often necessary to test transfer. But the tutor must know what changed.

A window can include different task forms if the target capability is stable and the variation is intentional. It should not merge unlike evidence simply because every result has a percentage sign.

The Task-Purpose Gate remains relevant: teaching work, practice, diagnosis, progress monitoring and verification do not mean the same thing.

Time does not make unlike evidence comparable.

Composite case: one disastrous lesson after six stable weeks

The following case is fictional and constructed for teaching.

Alicia has shown stable control of linear-equation solving across six weeks. She selects the method independently, manages signs well and succeeds on changed questions. One evening she makes repeated sign errors and needs two prompts.

Her tutor can tell at least three stories.

Story one: the old weakness has returned; reopen Repair immediately.

Story two: six weeks of strong evidence prove the bad lesson does not matter; ignore it.

Story three: the bad lesson is new evidence that deserves a bounded check before the route changes.

The tutor chooses the third.

The next session begins with two fresh equations under ordinary conditions. Alicia solves both accurately without prompts. A later mixed problem is also secure. The tutor records the poor lesson as a local disturbance rather than reopening the route.

The important point is not that a second good lesson always cancels a bad one. The point is proportionality. A major intervention should not be triggered by one discordant sample when the cost of waiting for a short confirmation is low and the prior evidence is strong.

If the error had created an immediate safety issue, high-stakes submission risk or imminent examination problem, the action threshold might be different.

The window belongs to the decision.

Composite case: the term average that refuses to notice the learner improved

This case is fictional.

Beatrice began the term weak in summary writing. Her first four pieces scored poorly. The tutor identified that she was copying details instead of selecting meaning. Over the next month, instruction changes. Three fresh summaries show a clear improvement in selection and compression under reduced support.

The term average remains mediocre because the four early pieces still count equally.

A parent asks whether Beatrice is improving. If the tutor reports only the average, the answer understates the current state. If the tutor reports only the last piece, the answer may overstate stability.

A better report separates history from current evidence.

"Her term average still includes the four early pieces from before the repair. In the last three comparable summaries, she has consistently selected the main ideas more accurately with reduced prompts. We are treating that as a promising new pattern, not yet as a permanent result. The next check is a fresh passage after a delay."

The review window has shifted because a targeted intervention created a plausible new learner state.

Older evidence has not been deleted. Its role has changed from current-state estimate to baseline history.

Use a short window for fast-changing decisions

Some tutoring decisions operate on a short timescale.

Did today's explanation land? Can the learner move from guided to independent practice? Is the current example too difficult? Did a new representation clarify the misconception?

For these questions, a short window can be appropriate because the decision itself is local and reversible.

The tutor may use the last two or three attempts, a short oral explanation and one changed item. Waiting for a month of evidence would be absurd.

The key is to keep the claim local.

"Ready for the next practice step today" is not the same as "mastered this topic".

A short window can support a short-horizon decision without pretending to prove long-term capability.

Use a wider window for stability claims

Other decisions require evidence over time.

Can a scaffold fade permanently? Has a repaired prerequisite remained available? Is the learner ready to leave special monitoring? Is examination performance becoming reliably stable?

These claims need a wider window because delay and changing conditions are part of the construct.

The Independence Test already asks whether unsupported success is stable across strategy selection, monitoring, recovery, help-seeking, transfer and delay.

The review-window contribution is simple: a stability claim cannot be supported by a window that contains only the immediate post-teaching period.

If durability matters, time must enter the evidence.

The window should reset after a real regime change

Sometimes the learner system changes enough that the old series should not be averaged with the new one as though nothing happened.

Examples include a major intervention, a change from typed to handwritten examination conditions, a new legitimate accommodation, a tutor reassignment with a different route, movement from one syllabus to another, or a long interruption followed by re-entry.

This does not mean historical evidence disappears.

It means the programme can mark a new comparison regime.

Before the change: useful baseline and history.

After the change: current-state evidence under the new condition.

The Accommodation Baseline already shows why access conditions must remain visible in comparison. The Modality-Shift Gate similarly warns that in-person and online evidence are not automatically interchangeable.

A regime marker protects the tutor from averaging across a structural break.

Do not choose the window after seeing which one tells the preferred story

The same learner record can produce several narratives depending on the window.

Last lesson: 45 per cent.

Last three: 63 per cent.

Last six: 71 per cent.

Term average: 68 per cent.

Year average: 74 per cent.

A tutor who selects the window after seeing which number supports the desired conclusion can unintentionally manipulate the evidence.

Where the decision is consequential, choose the window rule before looking at the result as far as practical.

For example: "We will review the last three comparable independent writing samples over at least three weeks, plus one fresh passage." Or: "We will decide whether to reopen this prerequisite after two fresh failures separated by a return to ordinary work, unless the error is high-cost enough to act sooner."

The Precommitted Route-Change Threshold owns the trigger rule. The window gate adds the evidence period from which the trigger is evaluated.

Precommitment limits hindsight bias.

Trend is not the same as sequence

Tutors often draw an imaginary line through scores.

62, 65, 68, 72. Upward trend.

But the sequence can be created by easier tasks, more prompting, increasing familiarity or a change in scoring standard.

The Criterion Drift Check protects the standard. The Adaptive-Difficulty Check protects interpretation when question difficulty changes.

A progress window should therefore contain condition notes, not just scores.

The tutor does not need a statistical model. One line can be enough:

"72%, fresh unseen set, no prompts, harder transfer mix."

That sentence is often more useful than another decimal place.

Composite case: progress appears flat because the tasks became harder

This case is fictional.

Ciara scores around 70 per cent for five consecutive Mathematics checks. The parent sees no progress.

The tutor reviews the window.

Check one contained direct routine questions with method labels. Check two removed labels. Check three mixed two nearby methods. Check four introduced changed representation. Check five added a realistic time limit.

The percentage stayed flat while the performance condition became more demanding.

It would be wrong to claim a precise amount of improvement. It would also be wrong to say nothing changed.

The tutor reports: "Her percentage is stable while the tasks have become less cued and more varied. That is evidence that the capability is travelling into harder conditions. We are keeping the score and the condition together rather than comparing percentages alone."

The temporal window becomes meaningful because it preserves how the task evolved.

A moving average can hide a turning point

Averages reduce noise. They also delay recognition.

Suppose a learner has ten low results followed by four strong results after a well-targeted repair. A twelve-sample average will still look mediocre for some time.

This is not a statistical defect. It is what an average does.

The tutor's job is to decide whether the recent cluster represents a plausible new state.

Useful questions include:

Did the improvement begin after a meaningful intervention?

Does it appear across fresh examples?

Did support reduce rather than increase?

Does it survive a delay?

Does it appear in a second source, such as school work or an independent paper?

If yes, the tutor can report a recent shift while keeping uncertainty visible.

"Recent evidence is stronger than the long-run average" is a legitimate statement.

"Therefore the learner is permanently fixed" is not.

One bad result can matter immediately when the error cost is asymmetric

A review window is not a rule against acting on one result.

The Error-Cost Asymmetry Gate exists because acting too early and waiting too long do not always carry equal cost.

If one result reveals that the learner cannot access a prerequisite needed for tomorrow's examination, waiting for three more samples may be irresponsible. If one response shows that a legitimate accommodation is missing, access should be repaired immediately. If a learner submits work that may violate authorship boundaries, the tutor should not average that concern with previous ordinary homework.

The window should widen when uncertainty is tolerable and narrow when the cost of delay is high.

That is not inconsistency. It is decision-sensitive evidence use.

Missing observations do not make the window neutral

A learner misses two sessions.

The review window now contains fewer observations and perhaps a selective sample of easier or more stable periods.

The Attendance Differential owns the interpretation of missed sessions. The Attrition Evidence Gap owns programme-level missing outcome data when learners disappear from final evidence.

At learner level, the review-window implication is simpler: absence reduces temporal coverage.

Do not say "stable over six weeks" when the learner was only observed twice.

State what the window actually contains.

The calendar can matter more than the number of samples

Three samples collected over one afternoon do not provide the same temporal evidence as three samples collected over three weeks.

If the decision concerns retention or repeated performance, spacing matters.

Likewise, six samples collected before a school holiday may say less about the learner's current state after a long interruption than two fresh post-holiday checks.

This is why review windows should sometimes be described by both sample count and elapsed time.

"Three fresh checks across four weeks" carries different information from "three checks today".

A tutoring record should not need sophisticated time-series analysis to preserve this distinction.

Parent reporting needs two clocks

Parents often want both the long view and the current view.

"How has my child done this term?"

"How are they doing now?"

These are different questions.

A good report can use two clocks.

The history clock summarises the broader period: where the learner began, what changed, what remains variable.

The current-state clock summarises the freshest comparable evidence relevant to the next decision.

For example:

"Across the term, accuracy has varied between 55 and 75 per cent because early papers exposed a prerequisite gap. In the last three fresh checks after repair, she has been above 70 per cent with fewer prompts. We are treating that recent cluster as the current working state, but we still need a delayed mixed check before calling it stable."

This avoids choosing between honesty about history and responsiveness to new evidence.

Review windows should differ by construct

Some capabilities can change quickly and be measured frequently. Others need longer observation.

A learner may correct a factual misconception within one lesson. Writing quality, self-regulation, examination pacing or independent study routines may require a longer period because the behaviour appears under varied conditions.

Do not impose one review cadence across every educational target.

A weekly check can be appropriate for one skill and meaningless for another.

The tutor should ask: how quickly could a real change plausibly appear, and how often can this capability be sampled without the measurement process dominating learning?

That is part of assessment validity.

Avoid the "latest three" superstition

Three is a convenient number.

It is not a law.

Three comparable independent attempts may be plenty for a narrow, low-noise decision. They may be far too little for a broad capability with high variation. Five observations can be redundant if they all come from the same worksheet family. Two observations can be compelling if they discriminate a very specific cause and are followed by a changed-condition check.

Use the smallest window that supports the decision with honest uncertainty.

Do not turn a practical heuristic into a pseudo-scientific threshold.

A practical window-building protocol

Start with the decision.

"What are we deciding now?"

Then define the evidence condition.

"What kind of performance would actually answer that question?"

Take the newest comparable sample and work backwards.

For each older sample, ask:

Does this describe the same target capability?

Was the support condition comparable or intentionally informative?

Has a major intervention or regime change happened since?

Does adding this sample improve the decision, or merely make the dataset larger?

Stop when older evidence becomes mainly historical context.

Then inspect the shape of the window.

Is one result dominating? Is the apparent trend created by changing task difficulty? Are all samples from the same source? Is there enough elapsed time for the claim being made? Is a missing interval being mistaken for stability?

Finally, write the conclusion at the same scale as the evidence.

"Current evidence supports moving to mixed practice."

"Recent improvement is promising but not yet durable."

"The latest decline needs a fresh check before route change."

"The old average no longer represents the current supported condition."

That language is far more useful than "the score is up".

Composite case: the route change that arrives one week too late

This case is fictional.

Denise has been struggling with a particular comprehension inference. Her term average is reasonable because earlier literal questions were strong.

Over four recent lessons, the same inference failure appears in fresh passages. Each time, the tutor records it but does not change the route because the overall average remains above the programme threshold.

By the fifth lesson, the pattern is undeniable.

The failure was not insufficient evidence. It was a review window that was too wide for the decision. Historical strength in other task types diluted a current repeated weakness.

A better window would have been target-specific and recent: the last four independent inference opportunities, not the entire comprehension average.

This is why a time window cannot be separated from construct coverage.

The unit being reviewed must match the job.

Progress review is an action system, not a reporting ceremony

The Stanford NSSA data guidance is useful because it connects assessment to instructional response. Data are not collected for admiration.

A review window should therefore end with a decision or a deliberate hold.

Continue the route.

Branch temporarily.

Collect one discriminating check.

Fade a scaffold.

Reopen a prerequisite.

Keep monitoring without change.

Escalate to a different owner.

If the review produces no possible action, the programme should question whether the evidence needs to be collected at that frequency.

The Evidence-Capture Burden Gate protects that principle.

The best review window is not the one with the most data. It is the one that supports a better next decision.

Failure modes

The single-lesson overreaction. One poor session triggers a major route change despite strong recent evidence and low cost of waiting for confirmation.

The historical-average drag. Old weak results continue dominating the learner model after a credible new pattern has emerged.

The cherry-picked window. The tutor chooses the time range after seeing which range supports the preferred story.

The mixed-condition trend. Scores collected with different support, difficulty or purpose are averaged as though they represent one stable measure.

The sample-count illusion. Several observations collected in one narrow period are treated as evidence of durability across time.

The stale-baseline problem. Evidence from before a major intervention, accommodation or curriculum change remains the default current estimate.

The latest-is-truth problem. The newest result receives full authority merely because it is newest.

The wrong-construct window. A broad subject average is used to hide a repeated failure in one consequential task family.

The no-action review. Tutors collect and summarise data that never change teaching, monitoring or communication.

The fixed-cadence superstition. Every skill is reviewed on the same timetable regardless of how quickly meaningful change could appear.

Recency weighting should remain interpretable

It can be tempting to solve the review-window problem with a formula.

Give the newest result 50 per cent weight, the previous result 30 per cent, and older evidence 20 per cent. The calculation looks disciplined.

Unless the weights are validated for the decision, they can create false precision.

The newest sample is not always more informative. A fresh but highly cued worksheet can deserve less authority than a slightly older independent examination-style response. A result after a major intervention may deserve separate interpretation rather than a larger numerical weight.

In ordinary tutoring, explicit qualitative weighting is often better.

"Most weight: last two independent fresh checks under current conditions."

"Supporting context: school paper from four weeks ago."

"Historical baseline: earlier supported work before the scaffold was removed."

This language tells another tutor why the evidence matters.

The goal is not to avoid numbers. It is to prevent an arbitrary formula from hiding professional judgement.

A disagreement inside the window is often the most useful part

Tutors naturally want a clean pattern.

But disagreement can reveal the boundary of capability.

The learner succeeds in short questions and fails in long ones.

They succeed on familiar representation and fail when the same relationship appears in a graph.

They write accurately at home and lose control under time.

They retrieve after one day and fail after two weeks.

Averaging these results can erase the condition that matters.

When the window contains systematic disagreement, partition before averaging.

"What separates the successful cases from the unsuccessful cases?"

That question can expose the next tutor decision more directly than the mean score.

The Progress-Review Window Gate therefore treats heterogeneity as evidence, not statistical inconvenience.

Close the window when the decision has been made

Review can become endless.

A tutor keeps collecting evidence because more evidence always feels safer. The learner remains in diagnostic limbo.

Once the predeclared decision threshold is met and the evidence window is adequate for the claim, act.

Move to the next route, hold the current route, fade support, reopen Repair or schedule the next delayed check.

Then begin a new window appropriate to the new decision.

This creates episodes of evidence rather than one continuously accumulating pile.

The record remains historical. The active decision gets a bounded evidence horizon.

Evidence boundaries and sources

The Stanford National Student Support Accelerator's current Section 4: Utilize Data recommends ongoing formative assessment, structured review of student data and longer-term progress tracking so tutors can personalise instruction and adjust in response to evidence. It also emphasises that tutors need time and support to analyse data. This is programme-design guidance, not a validated rule for a specific number of samples or weeks.

AERO's Monitor Progress, updated 14 May 2026, provides research-informed guidance on frequent checks for understanding and responsive instruction. It supports the feedback loop between evidence and teaching without prescribing a universal temporal review window for private tutoring.

The decision rules in this article are therefore professional assessment principles. No claim is made that "last three", "four weeks" or any other fixed window is universally correct.

The relevant test is whether the selected evidence remains comparable and current enough for the decision being made, and whether the conclusion stays proportionate to the window.

The end state

A learner record should not become a museum.

History matters. It tells the tutor where the learner has been, which repairs were needed, which conditions were fragile and which interventions have already been tried.

But tutoring is trying to change the learner.

The review system must therefore be capable of noticing a genuine new state.

The Progress-Review Window Gate protects both sides of that problem. It stops the tutor from chasing every fluctuation, and it stops the past from voting forever.

Start with the decision. Read the freshest comparable evidence. Widen the window until uncertainty becomes manageable. Mark structural breaks. Keep the conditions visible. Let older evidence become history when it no longer describes the current state.

Then make the next route decision from the window that actually belongs to now.