Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

The Tutor Handbook Vol No.0190 | The Implementation-Reach Gate — How a Tuition Programme Checks Whether a Teaching Practice Reaches the Learners and Tutors It Was Meant to Serve Instead of Succeeding Only in the Easiest Cases

The Tutor Handbook · Volume 0190 · Series ID THB-0190

The Tutor Handbook: Complete Series Index

A teaching practice can succeed perfectly where it appears and still fail the programme

A tuition programme introduces a new feedback routine. Tutors are trained. The routine looks strong in observed lessons. Learners revise work instead of merely reading comments. Follow-up checks show useful improvement.

The programme celebrates.

Months later, somebody asks a different question: which learners actually received the routine?

The answer is uncomfortable. Tutors used it mostly with established groups that finished tasks on time. Learners who frequently arrived with urgent school work rarely reached the feedback stage. Newly enrolled learners spent more time in diagnosis. Groups with greater reading support needs used shortened lessons and often skipped the routine. One tutor used it consistently; another used it only when the class was running smoothly.

The practice was feasible in some settings. It was often delivered with good fidelity where it was used. It may even have helped the learners who received it. But the learners for whom the programme claimed the routine was standard did not all have a meaningful opportunity to receive it.

That is an implementation-reach problem.

The Implementation-Reach Gate asks whether a teaching practice reaches the learners, tutors, groups and lesson conditions it was meant to serve—or whether the evidence of success comes mainly from the easiest places to implement it.

Reach is not the same as enrolment, attendance or popularity. It is the relationship between the intended opportunity and the opportunity that was actually delivered.

Quick answer

Before evaluating whether a teaching practice “works”, define who and what it is supposed to reach.

For each routine, identify the eligible learners, tutors, groups and lesson conditions. Then check whether those intended opportunities actually occurred. Do not count only successful uses. Include eligible occasions where the routine was skipped, shortened beyond recognition, displaced by another demand or made inaccessible by the delivery format.

When reach is uneven, investigate the pattern. Does the practice disappear in difficult groups, late sessions, online lessons, newer learners, examination periods or cases requiring access support? Is that absence justified because the practice is not appropriate there, or is the system quietly excluding the people it was meant to help?

Separate reach from fidelity. Reach asks whether the practice got there. Fidelity asks whether, once there, it was delivered with its intended mechanism intact.

Separate reach from outcome. A practice can show excellent outcomes among recipients while systematically missing part of the intended population.

The programme should improve the route to the practice before generalising from the easiest recipients.

AERO’s new implementation guide makes reach a named outcome

The Australian Education Research Organisation published Staying on track: Monitoring implementation outcomes on 15 September 2026. The guide is designed for school implementation teams and is research-informed professional guidance rather than a private-tutoring trial.

It distinguishes implementation outcomes including feasibility, acceptability, fidelity, reach and sustainability. That distinction is useful because education programmes routinely jump from “we introduced the practice” to “the practice is working” without checking whether the intended people actually received the practice in a meaningful form.

The earlier Education Endowment Foundation implementation guidance, third edition published 24 April 2024, similarly treats implementation as a structured process shaped by context, people and supporting conditions rather than a one-time launch.

For tutoring, the reach question deserves its own owner because small-group programmes can look highly personalised while still producing selective implementation. A routine may reach one learner in a three-learner group and not the other two. It may reach confident tutors and not tutors who need the most support to use it. It may reach stable Alignment work and disappear from messy Repair cases.

Those patterns matter before outcome claims are made.

Reach begins with a denominator

A programme cannot know whether reach is adequate until it knows who was supposed to receive the practice.

Suppose a tutor uses a retrieval routine in twelve lessons. That number sounds substantial. But how many lessons were eligible? Twelve out of fourteen tells one story. Twelve out of sixty tells another.

The same logic applies at learner level. Twenty learners used a new self-explanation routine. Was the routine intended for those twenty only, or for ninety learners who had reached the relevant knowledge threshold?

The denominator should not be invented after the result is known. Before implementation, declare the intended opportunity as clearly as practical.

“Use this routine once each week with learners who have established the target knowledge and are now practising method selection.”

“Use this feedback action loop when the learner receives substantive corrective feedback and has enough lesson time for a reattempt.”

“Use this three-learner comparison routine when the group shares a common target and each learner can make a private first response.”

These eligibility conditions make the denominator educational rather than merely administrative.

Universal reach is not always the goal

A reach gate can be misused if programmes assume every practice should touch every learner.

Some routines are inappropriate for some conditions. A sophisticated self-explanation prompt may overload a learner who is still acquiring basic knowledge. A three-learner peer-comparison routine does not belong in a one-to-one session. A timed performance routine may be inappropriate during early Repair. A digital tool may not be necessary when the same learning job is better served offline.

So reach must be evaluated against intended eligibility, not total enrolment.

The question is not, “Did everybody get it?”

It is, “Did everybody who should reasonably have had this learning opportunity actually receive it, and did we exclude anyone for reasons that belong to the system rather than the educational purpose?”

This distinction prevents implementation monitoring from turning into compulsory uniformity.

The easiest cases can dominate the evidence

Successful implementations often become visible first in favourable conditions.

Tutors try a new routine with learners they know well. Leaders observe groups with strong attendance. Digital pilots start with families who already have reliable devices. New coaching moves are demonstrated by the most experienced tutor. Data are easiest to collect from learners who finish tasks within the planned time.

Those choices are understandable during development. They become misleading when the resulting evidence is presented as representative of the entire programme.

The core reach question is therefore: who is missing from the evidence?

If the routine is absent mostly from learners whose needs are more complex, then the programme may have learned that the routine works where implementation is easiest, not where its promised value is greatest.

This does not automatically invalidate the practice. It tells the programme what implementation problem remains unsolved.

Reach is different from feasibility

The Implementation-Feasibility Gate asks whether ordinary tutors can realistically carry a routine within available time, materials, skills and operating conditions.

Reach asks whether the practice actually arrives at the eligible opportunity.

A routine can be feasible but have poor reach because tutors forget it, scheduling excludes certain groups, materials are distributed unevenly or leaders never clarify eligibility.

A routine can also have broad reach but be infeasible in the long run: tutors perform it everywhere for one month by working extra unpaid hours, then the system collapses.

Separating these diagnoses matters. A reach problem may require better workflow, clearer eligibility, distribution or monitoring. A feasibility problem may require redesign, more capacity or narrower scope.

Calling both “low implementation” loses the decision.

Reach is different from fidelity

The existing Implementation Fidelity Check asks whether the intended active ingredients, dose, sequence and support conditions were actually delivered.

Reach asks whether the intended learner or tutor got a real opportunity for the practice in the first place.

Consider four groups eligible for a new retrieval routine.

Group A receives the routine exactly as intended. Group B receives a shortened version that removes corrective feedback. Group C never receives it because urgent school work repeatedly displaces the opening review. Group D receives it only from a covering tutor once.

Group A has reach and strong fidelity. Group B has reach but questionable fidelity. Group C has a reach problem. Group D has partial or intermittent reach.

These distinctions lead to different actions.

Without them, the programme might average everything into a vague “implementation rate”.

Reach is different from attendance

A learner can attend every lesson and never receive the intended practice.

Attendance tells the programme whether the learner was present. It does not tell the programme what learning opportunities occurred.

The reverse is also possible. A routine may have excellent reach among attending learners, but low attendance means the learner receives little total exposure. That becomes a dosage or attendance problem rather than an implementation-reach problem inside attended lessons.

The Attendance Differential already asks what missed sessions mean before treating them as a motivation problem. Reach should not absorb that job.

The useful separation is simple:

Was the learner present?

If present and eligible, did the intended opportunity occur?

If it occurred, was it implemented well enough to count as the intended practice?

If it was implemented, what did the learner do with it?

Each question owns a different part of the chain.

Composite case: feedback reaches only the fast finishers

The following case is fictional and constructed for teaching.

A three-learner English group uses a feedback-action routine. Learners draft a paragraph, receive one high-value feedback point, revise, and later attempt a fresh paragraph.

Alicia writes quickly. She usually reaches the revision stage. Beatrice needs more planning time. Ciara receives legitimate language support and often completes the first paragraph near the end of the lesson.

After six weeks, Alicia has used the full feedback-action loop five times. Beatrice has used it twice. Ciara has never used it. The tutor has technically “implemented the feedback routine” in the group.

The reach gate reveals a distribution problem.

The programme can respond in several ways. It might shorten the initial writing sample so all learners reach the feedback stage. It might start feedback on work begun previously. It might schedule the routine across two lessons. It might use separate but equivalent targets for learners whose writing rate differs.

The wrong response is to celebrate Alicia’s improvement as evidence that the routine is embedded across the group.

The practice has reached the learner whose lesson conditions make it easiest to deliver.

Reach can fail inside a three-learner group

Small groups tempt leaders to think reach is automatic. There are only three learners. Surely everybody gets the routine.

Not necessarily.

One learner answers first and receives most follow-up questions. One learner is regularly pulled into a prerequisite branch while the other two complete the shared routine. One learner misses the crucial comparison because they are finishing a previous task. One learner’s access support is not compatible with the digital format, so the tutor substitutes generic independent work.

The class has reach at group level and unequal reach at learner level.

This is another reason group-level implementation records should not automatically stand in for individual opportunity when the practice is intended to affect individuals.

The Individual Accountability Gate owns the evidence problem inside collaboration. Reach adds the implementation question: did each eligible learner actually get the learning opportunity the programme intended?

Tutors themselves can be the intended recipients

Reach is not only about learner-facing practices.

A coaching programme can be designed for all tutors and reach only those whose schedules align with the coach. New tutor training can be available in principle but inaccessible to part-time tutors who teach only evenings. A material-update briefing can reach the tutors who attend a meeting and miss tutors on leave. A new observation routine can be used with tutors in one location but rarely with those working online.

If the programme then compares teaching quality across tutors, unequal professional-learning reach becomes part of the interpretation.

This matters particularly when programme leaders say, “All tutors were trained.” Attendance at one training event is not automatically meaningful reach if the practice requires rehearsal, feedback or live coaching that some tutors never received.

The Coaching-Support Taper Gate and Observation-Sample Gate help interpret professional support. Reach asks whether intended support entered the tutor’s actual work at all.

Composite case: a dialogic routine disappears in difficult groups

This case is fictional.

A programme trains tutors to use a three-step dialogic routine: private thinking, learner explanation, peer comparison.

In stable groups, tutors use it often. In groups with frequent off-task behaviour or divergent needs, tutors revert to direct explanation and worksheets because discussion feels risky.

Leaders observe several successful dialogic lessons and conclude the routine is working.

A reach audit shows that the routine is systematically absent from the groups in which tutors have the least confidence facilitating learner talk.

There are several possible interpretations. Perhaps the routine truly is inappropriate in those conditions. Perhaps it needs stronger preconditions. Perhaps tutors need more facilitation training. Perhaps material design needs to reduce the risk of answer leakage. Perhaps the programme is unintentionally withholding a valuable learning opportunity from learners whose groups are harder to manage.

The reach gate does not decide which interpretation is correct. It makes the selective pattern visible so the programme can investigate rather than generalise from favourable cases.

The denominator should track opportunity, not paperwork

A programme can build a reach dashboard that is technically neat and educationally useless.

If tutors simply tick “used routine: yes/no” after every lesson, the data can become compliance theatre. Tutors may interpret eligibility differently. A five-second mention can count the same as a full learner opportunity. Missing data may be treated as non-use even when the lesson was not eligible.

A better reach record is minimal but explicit.

Was this lesson or learner eligible for the routine?

If yes, did the meaningful opportunity occur?

If no, what broad reason explains the absence: insufficient time, competing priority, learner not ready, material/access barrier, tutor choice, technical failure, or other?

The programme does not need detailed narrative for every lesson. It needs enough information to see systematic exclusion.

If almost every “not used” reason is “time”, the problem may be feasibility. If non-use clusters in one tutor’s groups, coaching or interpretation may matter. If non-use clusters in one learner population, access or design needs investigation.

Reach monitoring should create decisions, not merely percentages.

Beware the numerator without the denominator

“Eighty learners used the new tool” sounds impressive.

Without knowing how many eligible learners there were, the number is almost meaningless.

“Tutors completed 150 feedback loops this term” has the same problem. Were there 170 meaningful opportunities or 700?

Programmes naturally report positive counts because positive counts are easy to celebrate. Reach requires the less glamorous denominator.

This is especially important for AI and digital tools, where logins, messages and minutes can create an illusion of broad implementation. The Human-Supported AI Engagement Gate already warns that more screen time is not automatically more learning opportunity.

Reach adds: even if the tool is educationally useful for users, who never receives the intended supported use?

Eligibility should be documented before exclusion becomes habit

When a tutor decides a learner is “not ready” for a routine, that may be professionally sound. It can also become self-perpetuating.

A learner is excluded from self-explanation because knowledge is weak. Because they are excluded, they receive fewer opportunities to practise explanation. Months later, weak explanation is treated as evidence they were never suitable.

The solution is not forced participation. It is an eligibility review.

What prerequisite is missing? What evidence would show readiness? Is there a simpler version that preserves the learning job? When will the decision be revisited?

This keeps temporary non-reach from turning into permanent opportunity denial.

The same principle applies to tutor coaching. A tutor may not yet be ready for autonomous use of a complex routine, but the programme should have a route toward readiness rather than simply keeping the tutor outside the practice indefinitely.

Reach has an equity dimension without requiring identical treatment

Educational equity does not mean every learner receives identical routines.

It does mean the programme should notice if barriers such as schedule, delivery mode, language, disability access, group composition or perceived behaviour systematically reduce access to a valuable eligible opportunity.

Suppose an online homework-feedback tool cannot be used effectively with screen-reader software. The programme may say the routine is available to everyone because every learner has an account. Operationally, the opportunity is unequal.

Suppose the richest discussion tasks are reserved for the “top group” because tutors believe weaker learners need worksheets first. That may reflect a legitimate sequence in some cases. It may also become a low-expectation loop if learners never receive the knowledge-building support that would make rich discussion accessible.

The reach gate asks for evidence at the boundary rather than assuming fairness from intent.

Reach can vary by time of day and calendar

Implementation patterns often follow the timetable.

Early sessions may use the full routine because tutors are fresh. Late sessions may be shortened. Examination weeks may crowd out cumulative review. Monday groups receive a material update immediately; Saturday groups continue using the old version for days. Tutors with long transitions between sites may skip documentation that tutors in one location complete reliably.

None of these patterns automatically means wrongdoing.

They do mean that “programme-wide implementation” can hide time-based reach differences.

A useful audit occasionally cuts the data by session type, tutor, learner level, modality or calendar period to see whether absence clusters somewhere predictable.

The purpose is not surveillance. It is to find implementation geometry.

Composite case: a progress check reaches only the learners who look stable

This case is fictional.

A programme introduces a short monthly progress check intended for every learner in Alignment mode. Tutors use it consistently with learners who appear stable. When a learner has been struggling, tutors often skip the check and spend the time reteaching instead.

This sounds sensible. Why test a learner who obviously needs help?

The problem is that the programme now has systematic progress-check evidence only from the learners who look successful. Struggling learners receive more teaching but less comparable evidence about whether the intervention is changing the target capability.

The solution is not to force the full check at an inappropriate moment. The programme can create a smaller eligible sample or schedule the check after a defined repair window. It can explicitly say that active Repair learners are temporarily outside the Alignment monitoring routine and require a different evidence route.

What matters is that the absence is declared rather than invisible.

Otherwise, programme progress data become selected by perceived success before the measurement occurs.

Reach should be checked before outcome comparison

Suppose one tutor’s learners show stronger average improvement after a new routine is introduced.

Before crediting the routine or the tutor, ask whether exposure differed. Perhaps that tutor used the practice with nearly every eligible learner while another used it only with half. Perhaps the second tutor’s groups had more interrupted attendance. Perhaps access conditions prevented full use in one modality.

Outcome interpretation without reach information can confuse intervention effect with implementation distribution.

The Tutor-Effect Variation Gate already protects against ranking tutors from raw score gains without considering case mix and other conditions.

Reach adds one more implementation condition: did learners actually receive the practice being discussed?

A weak outcome after low reach does not tell the same story as a weak outcome after high-fidelity, high-reach implementation.

Reach can be partial rather than binary

Some learning opportunities occur incompletely.

A feedback routine may include learner reattempt but not delayed follow-up. A coaching routine may include observation but no debrief. A comparison routine may occur only for one of three learners. A retrieval review may happen two weeks out of four.

Treating reach as yes/no can hide this.

The programme does not necessarily need a complicated scale. It can distinguish meaningful categories such as expected exposure, intermittent exposure, minimal exposure and no exposure when those distinctions change decisions.

The important rule is to avoid making the categories more precise than the records support.

If the programme cannot distinguish a two-minute mention from a full learning opportunity, it should not pretend it can.

The implementation-reach receipt should stay small

A small tuition centre should not create a research bureaucracy around every routine.

The reach receipt should answer only enough questions to act.

Who was meant to receive the practice?

Who actually received a meaningful opportunity?

Where is non-reach clustering?

Are those exclusions educationally justified, temporary and reviewed—or are they symptoms of feasibility, access, training or workflow problems?

Once the programme understands the pattern, it can stop measuring if the measure no longer improves decisions.

Implementation monitoring should itself pass a feasibility gate.

Parents and learners can reveal hidden non-reach

Families sometimes expose implementation differences before dashboards do.

One parent says, “My child always gets a chance to revise after feedback.” Another says, “The comments are written, but the lesson moves on.” One learner describes regular mixed practice. Another in the same level says they have never seen it.

These reports are context, not proof. The Family-Evidence Integration Gate keeps parent observations separate from direct learner evidence.

But repeated reports can trigger a reach check.

The programme should ask whether the difference reflects legitimate route variation or accidental implementation inequality.

That is a much better response than promising that “all classes follow the same system” when the evidence has not been checked.

Reach and sustainability interact

A practice can begin with broad reach and narrow over time.

Initial training is fresh. Leaders remind tutors. Materials are stocked. Six months later, new staff have joined, the shared template has drifted, one device has failed, and examination pressure has squeezed the routine from late lessons.

A sustainable practice needs mechanisms that preserve reach as conditions change.

This does not mean permanent monitoring. It means occasional sampling after turnover, material changes, modality shifts or major programme growth.

If the routine slowly survives only in the strongest tutor’s classes, the programme has not sustained programme-wide reach even if that tutor still implements it beautifully.

Failure modes

The enrolment-equals-reach failure. A learner is counted as having access because they are enrolled in the programme, regardless of whether the practice occurs.

The numerator-only failure. Leaders celebrate the number of uses without knowing the eligible denominator.

The easiest-case failure. Evidence comes mainly from stable groups, confident tutors or learners who finish on time.

The attendance-confusion failure. Presence in the lesson is treated as evidence that the intended opportunity occurred.

The fidelity-confusion failure. A practice is implemented beautifully where it appears, so leaders assume it appears everywhere it should.

The universalism failure. Reach is maximised by forcing a routine into conditions where it is educationally inappropriate.

The permanent-not-ready failure. Learners or tutors excluded for legitimate short-term reasons are never reviewed for later eligibility.

The access-by-account failure. Digital availability is counted as reach even when the format is not usable under legitimate access conditions.

The group-level masking failure. A routine is marked present because it occurred in the group, although one learner never received the relevant opportunity.

The measurement-burden failure. Reach monitoring becomes so elaborate that it creates a new implementation problem.

A practical reach audit

Choose one important routine, not twenty.

Define the eligible opportunity in plain language. Sample a bounded period. For each eligible case, record whether a meaningful opportunity occurred. Look for systematic absence by tutor, group, modality, learner condition or timetable period.

Then investigate one pattern at a time.

If late classes have lower reach, is time the constraint? If new learners are excluded, is the routine genuinely inappropriate during diagnosis or has nobody defined an entry condition? If one tutor rarely uses the routine, is training, belief, material fit or workload the issue? If access-supported learners receive a different version, does that version preserve the same learning job?

Repair the route rather than chasing a percentage.

Then sample again later. If reach improves without fidelity collapsing, the implementation system is stronger.

Evidence boundaries and sources

AERO’s Staying on track: Monitoring implementation outcomes, published and updated 15 September 2026, is the central current implementation source for the reach concept used here. It is research-informed school guidance and not a validation study of private tutoring.

The Education Endowment Foundation’s Implementation guidance, third edition published 24 April 2024, supports examining implementation processes, context and supporting structures rather than assuming adoption equals delivery.

The National Student Support Accelerator’s current Toolkit for Tutoring Programs, Session Content guidance and Model Dimensions guidance provide tutoring-specific design context around grouping, tutor type, ratio, delivery and instructional need. They do not provide a universal implementation-reach threshold.

The reach categories and practical audit proposed in this article are professional design tools. They should remain proportionate to the decisions a tuition programme actually needs to make.

The end state

A programme should not be satisfied because a practice works beautifully somewhere.

It should know whether the practice reaches the learners and tutors for whom it was designed, under the conditions in which the programme promises it will operate.

If a routine disappears whenever the group is difficult, whenever time is tight, whenever accessibility needs change the format or whenever a less confident tutor is teaching, that pattern is part of the intervention story.

The answer is not to force universal use. It is to declare eligibility, make opportunity visible, distinguish justified exclusion from system failure, and repair the route where the practice is supposed to travel.

That is the Implementation-Reach Gate.

Before asking whether the routine worked, ask who actually had a fair chance to receive it.

Reach should be checked at the point where opportunities diverge

One final way to keep reach monitoring useful is to inspect the moment where eligible learners stop receiving the same opportunity.

In many programmes, the divergence is not at enrolment. It happens later. One learner needs longer diagnosis, so the shared practice is postponed. One tutor loses five minutes to a school-homework interruption. One online learner cannot use the planned tool. One group’s material arrives late. One learner finishes slowly and the reattempt falls beyond the lesson boundary.

Those points of divergence are more actionable than an end-of-term percentage because they reveal the mechanism of non-reach.

If the practice repeatedly disappears after a particular transition, redesign that transition. The solution might be a shorter entry task, a second-session handoff, an accessible format, an offline alternative, a clearer eligibility rule or a different lesson sequence. The exact response depends on what the missing opportunity was supposed to accomplish.

This also protects tutors from unfair judgement. A tutor who skips a routine because the learner is not yet eligible is making a different decision from a tutor who skips it because the materials were not prepared. A dashboard that records only “not used” erases that difference. Reach monitoring should preserve enough reason to distinguish justified adaptation from preventable absence without requiring a narrative essay after every lesson.

The same principle helps when a programme expands. New locations, new tutors and new delivery modes create fresh divergence points. A routine that had near-complete reach in one small centre may narrow when materials, coaching or scheduling become less direct. Rather than assuming growth preserves implementation automatically, sample the new edges of the system.

Reach is therefore not a one-time audit. It is a periodic question asked when the route changes: where should this practice now be able to travel, and what is stopping it from arriving there?