Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

The Tutor Handbook Vol No.0071 | The Applicability Check — How a Tutor Decides Whether Research From Another Age, Subject, Country or Setting Belongs in This Learner’s Route

The Tutor Handbook · Volume 0071 · Series ID THB-0071

The Tutor Handbook: Complete series index.

A tutor reads a strong study.

The study is carefully designed. The finding is interesting. The result appears relevant to learning.

There is one problem.

The participants are university students. The tutor teaches Primary 6.

Or the study is about vocabulary, while the tutor wants to use the idea in Additional Mathematics.

Or the intervention ran four times a week with trained staff inside schools, while the tutor sees a learner for ninety minutes once a week.

Or the research was conducted in another country, with a different curriculum, language environment, school structure and assessment system.

Or the study measured immediate recall, while the tutor cares about transfer three weeks later under independent examination conditions.

Does the research still matter?

Sometimes yes. Sometimes partly. Sometimes only as a mechanism clue. Sometimes not enough to justify changing the learner’s route.

The Applicability Check is the tutor’s disciplined test of how far research evidence can travel from the people, tasks, settings, implementation conditions and outcomes actually studied to the learner and tutoring decision in front of us.

The purpose is not to demand a Singapore private-tuition randomised trial for every teaching decision. That standard would make useful evidence nearly impossible to use. The purpose is also not to say, “Research shows this works,” whenever a paper uses a familiar educational word.

Good tutors need a middle discipline: respect strong evidence, inspect its boundaries, understand the proposed mechanism, compare the research context with the learner context, reduce confidence when important differences grow, and use bounded reversible implementation when direct evidence is incomplete.

Quick answer

Before changing tutoring because of research, state the research claim precisely. Identify who was studied, what they did, what the comparison was, how the approach was implemented, what outcome was measured and when it was measured. Then map those features against the learner’s age, prior knowledge, subject, language, setting, tutoring dose, support conditions, school demands and the outcome you actually care about.

Next ask whether the mechanism that supposedly produced the result is likely to operate here. A study about retrieval may offer a plausible mechanism across subjects, but its exact effect size may not travel from university prose recall to Primary Science explanation. A study of small-group tutoring may support the value of intensive targeted instruction, but it does not automatically validate a particular three-student private-tuition model with different frequency, tutor preparation and school integration.

When evidence is strong but context fit is uncertain, do not discard it. Downshift the claim. Use the evidence to design a cautious trial, define the learner receipt, preserve the known-good route, and stop or revise if the predicted mechanism does not appear.

1. What this volume owns

This volume owns the tutor’s decision about whether research is applicable enough to inform a specific learner route. It does not own the broader question of how scientific or educational research is conducted. It does not replace systematic reviews, research appraisal, statistical inference or formal evidence standards. It does not claim that tutors can determine external validity perfectly from a checklist.

AERO’s current Assessing Research Evidence guide makes the underlying professional problem explicit: evidence needs to be assessed for both reliability and relevance to context. Its companion Applying Research Evidence guide states that even after rigorous and relevant evidence is identified, educators still need to decide whether and how to apply an approach in their own learning environment.

The Institute of Education Sciences’ current Standards for Excellence in Education Research likewise treats generalizability as an explicit research concern: studies should define populations of interest, support valid estimates for relevant groups, report sample characteristics and consider implementation and scale. The What Works Clearinghouse makes contextual study information available precisely so users can examine “what works for whom and under what conditions”.

The Applicability Check translates that problem into a tutor-facing question: what, exactly, is this research evidence permission to do with this learner?

2. Strong research can still be a poor fit

Quality and applicability are related but different.

A well-designed study can estimate an effect credibly for the studied population and still leave uncertainty about another population. A weaker study can happen to resemble your learner closely and still provide poor causal evidence. Tutors need both questions:

  • Can I trust the research claim in the context that was studied?
  • How far can I transport that claim into my context?

Do not compensate for weak evidence by saying the context looks similar. And do not reject strong evidence merely because the country, subject or age differs. The first task is to understand the finding honestly. The second is to inspect the distance it must travel.

This distinction prevents two common errors. Evidence enthusiasts sometimes overgeneralise: “Retrieval practice works, therefore every lesson should start with a closed-book quiz.” Evidence sceptics sometimes overlocalise: “That study was not done in Singapore, so it tells us nothing.” Both positions avoid the harder work of mechanism and context.

3. Begin with the actual research claim

Research often reaches tutors after several rounds of compression.

A paper becomes a university press release. The press release becomes a professional-development slide. The slide becomes “spacing is better”, “feedback should be immediate”, “interleaving works”, “AI improves learning” or “small groups are effective”.

Before asking whether the claim applies, reconstruct it.

Who participated? What task did they perform? What exactly did the intervention change? What did the comparison group do? How long did the intervention last? Who delivered it? How much training did they receive? What outcome was measured? Was the measure immediate, delayed, near transfer, far transfer, classroom attainment, self-report or something else? How large and uncertain was the effect? Was the study designed to estimate causation or only association?

A tutor does not need to become a research methodologist to ask these questions. They are basic safeguards against applying a headline instead of applying evidence.

4. Population: who was studied?

Age matters, but age is not the only population feature.

Prior knowledge can change how an instructional strategy behaves. A worked example may support a novice differently from a learner who already has a strong schema. Language proficiency can alter the demands of a task. Students selected because they are substantially behind grade level may respond differently from students already performing strongly. Volunteers in an online study may differ from students receiving compulsory school intervention.

The IES guide Enhancing the Generalizability of Impact Studies in Education focuses on this problem from the research side: define the target population, recruit samples deliberately, assess differences between sample and target, and report enough information for generalizability to be examined.

The tutor uses the same logic in reverse. Who were these learners, and which characteristics could plausibly change the mechanism? “Different country” is too broad. “Participants were adult university students with strong reading fluency, while my learner is a ten-year-old still building the vocabulary needed to understand the task” is an actionable difference.

5. Subject: does the mechanism survive the knowledge structure?

Some learning mechanisms plausibly travel across subjects. Retrieval requires reconstructing information without the answer present. Spacing changes the timing of encounters. Feedback provides information about a performance. Comparison can reveal structure.

But subjects place different demands on those mechanisms.

Retrieving a factual definition is not the same as constructing an argument. Interleaving several Mathematics problem types trains discrimination among procedures in a way that may not map neatly onto mixing unrelated English writing genres. A Science explanation requires causal relationships and model use; a vocabulary task may primarily require form-meaning mapping and contextual selection.

When transferring evidence across subjects, keep the mechanism and re-earn the application. Ask: what is the target operation here? Does the research strategy actually require that operation? What subject knowledge must already exist? What neighbouring demands could change the effect?

A useful general principle can survive while the exact practice design changes substantially.

6. Setting: where did the intervention live?

A strategy delivered inside a daily school timetable may depend on conditions a weekly tutor does not have.

Consider high-impact tutoring research. Programmes may operate several times a week, coordinate closely with classroom teachers, use common curriculum materials, employ trained tutors and collect structured progress data. The National Student Support Accelerator’s Tutoring Quality Standards distinguishes recommendations grounded in robust research from research-informed and emergent recommendations precisely because implementation features differ in evidential strength.

A private tutor teaching three students once a week should not say, “Research on high-impact tutoring proves this model.” The frequency, programme structure, tutor preparation, school integration, learner population and measurement may be materially different.

That does not make the high-impact tutoring research irrelevant. It may identify useful design features: consistency of tutor, targeted instruction, alignment, use of data, sufficient dosage, tutor support. The Applicability Check converts those features into questions rather than inherited claims.

7. Dose: how much of the intervention was actually delivered?

An intervention is not merely a name.

“Tutoring”, “retrieval practice”, “coaching”, “peer feedback” and “worked examples” can describe programmes with very different frequency, duration, spacing, quality and support.

If a study delivered twenty minutes of targeted practice four times a week, reproducing the label once every Saturday may not reproduce the mechanism. If professional coaching included observation, feedback and repeated rehearsal, a single introductory workshop is not the same intervention. If a scaffold was faded over multiple lessons after stable performance, removing it immediately is not the same procedure.

The tutor should identify the active ingredients and the dose they plausibly require. Then ask what can be implemented faithfully within the actual tuition schedule.

Where the local dose is much smaller, do not assume a proportionally smaller version of the same effect. Some mechanisms may need a threshold of practice or time; others may function differently when compressed. Mark that uncertainty.

8. Provider: who delivered it, and with what preparation?

Educational interventions are often human-delivered.

A study may involve trained teachers, researchers, specialist coaches or tutors receiving manuals and ongoing support. The effectiveness of the named strategy may partly depend on judgement, feedback quality, relationship, subject knowledge or implementation fidelity.

This matters for tutor application. Reading that formative assessment is valuable does not automatically train a tutor to ask discriminating questions. Reading that scaffolding should fade does not teach the tutor when the learner is ready. Reading that feedback can improve learning does not prevent feedback from becoming answer substitution.

Volume 0070, The Rehearsal-to-Live Gate, addresses this implementation layer directly. Applicability depends not only on whether the mechanism can travel, but whether the tutor can enact it under local conditions.

9. Outcome: what did the study actually improve?

The outcome is often where overgeneralisation becomes invisible.

A study shows better immediate recall. The tutor claims better examination performance. A programme improves a standardised reading score. The tutor claims greater learner independence. A tool improves the quality of submitted essays. The tutor claims the learner became a better writer. A short intervention improves accuracy on practised problem types. The tutor claims transfer.

Those extensions may eventually be true. The original outcome does not establish them.

Write the outcome precisely: immediate factual recall, delayed retention after one week, method-selection accuracy, essay quality with tool access, self-reported confidence, attendance, standardised attainment, time to completion. Then compare it with the outcome needed in the learner route.

If the tutor cares about independent transfer, a study measuring supported immediate performance may be mechanism-relevant but outcome-distant. That distance should reduce confidence, not erase interest.

10. Time: when was the outcome measured?

Learning is time-sensitive.

An intervention can improve performance during or immediately after practice while producing a different pattern after delay. Familiarity, working memory of feedback, temporary support and recent repetition can all affect short-term results.

For tutoring, ask whether the research outcome matches the time horizon you need. If the study measured performance five minutes later and the learner must retrieve the idea three weeks later, the research may still support a mechanism but does not directly answer the retention question.

This is why the eduKateSengkang learning route repeatedly returns to delayed checks and changed conditions. The local implementation should produce the receipt that matters locally, even when the research measured something nearer.

11. Mechanism can travel farther than an effect size

This is one of the most useful distinctions in evidence application.

Suppose a study finds that one form of retrieval practice improves delayed memory by a particular amount in a particular sample. The exact numerical effect depends on study design, material, baseline, comparison condition, retention interval and population.

The underlying mechanism—that attempting to retrieve previously learned material creates a different learning event from simply rereading it—may plausibly travel much farther than the exact effect estimate.

A tutor can therefore use the research to justify testing retrieval as a practice condition without promising the same magnitude of benefit. The local question becomes: does this learner retrieve more reliably after delay, under the subject conditions we care about, without creating unacceptable cost?

This distinction is useful beyond retrieval. A mechanism can inspire an intervention while the original study’s effect size remains local to its design. Do not turn a plausible mechanism into a universal percentage.

12. Context distance should change the strength of the claim

Think of applicability as a gradient rather than a binary verdict.

When population, task, setting, implementation and outcome are very similar, the evidence may justify a relatively confident expectation—subject to the quality of the research itself. When several dimensions differ but the mechanism is plausible, the evidence may justify a research-informed trial rather than a strong prediction. When the only connection is a broad analogy, the evidence may suggest a question but not an intervention claim.

The National Student Support Accelerator’s current quality standards offer useful language here. Recommendations can be research-based, research-informed or emergent depending on the evidence supporting them. A tutor does not need to copy that classification system into every lesson, but the underlying discipline is powerful: say how direct the evidence is.

“Research proves this will work for your child” is usually too strong. “There is strong evidence for this learning principle, but the exact implementation differs from our setting, so I am using it to design a bounded trial and checking whether the predicted improvement appears” is more honest and more useful.

13. The Applicability Map

A tutor can map research to a learner across seven practical dimensions.

  • People: age, prior knowledge, language, learner profile and selection criteria.
  • Knowledge: subject, topic, task structure and prerequisite demands.
  • Place: school, tuition, home, online or laboratory setting and the surrounding institutional conditions.
  • Provision: frequency, duration, group size, tutor preparation, materials, supports and implementation fidelity.
  • Performance: what learners were actually asked to do.
  • Proof: which outcome was measured, under what support conditions and after what delay.
  • Pathway: the mechanism that is supposed to connect the intervention to the outcome.

The alliteration is only a memory aid. It is not a validated evidence score. Do not assign points and pretend the sum quantifies generalizability.

The map helps the tutor ask where the important distances are. A country difference may be superficial if the task and mechanism are highly similar. A tiny difference in support condition may be critical if independence is the target. Context distance is causal, not geographical.

14. Country differences: do not overreact in either direction

“This was done in the United States” is not an automatic rejection criterion.

Basic memory, attention and practice mechanisms do not change at national borders. Many instructional principles can be informative across systems.

But educational institutions do differ. Curriculum sequence, examination format, class size, language environment, teacher roles, school calendar, available technology, tutoring frequency and cultural expectations can change implementation and outcome relevance.

A Singapore tutor should therefore separate mechanism from system claim. Research from another country may inform how worked examples or retrieval function, while Singapore-specific syllabus and examination decisions should use current local owners and official MOE/SEAB information when relevant.

Do not write “international research proves this is best for PSLE” unless the evidence genuinely supports that specific claim. Usually it does not.

15. Age differences: development and prior knowledge both matter

Age is often used as shorthand for many things at once: reading fluency, metacognitive skill, working-memory strategies, self-regulation, vocabulary, subject knowledge and independence.

When transferring a study from adults to children, ask what part of the effect depends on capacities the younger learner may not yet have. A university student can follow a complex self-directed study protocol with little coaching. A Primary learner may need the same underlying strategy embedded in a simpler routine with more modelling and external structure.

Conversely, do not assume children need more help merely because they are younger. A well-practised ten-year-old may manage a familiar retrieval routine more independently than an adult novice facing an unfamiliar subject.

Age matters through mechanism and experience. Ask what the learner must understand, remember and regulate for the intervention to function.

16. Research-to-tuition case: retrieval practice and PSLE Science

Suppose a tutor reads strong research showing benefits of retrieval practice for delayed memory in educational tasks. The temptation is to create daily closed-book Science quizzes and call the method evidence-based.

The Applicability Check slows the move down.

What is the local target? If Ciara is forgetting definitions and causal relationships she has already understood, retrieval may be a close mechanism match. If she has never understood the mechanism, retrieval cannot replace instruction. If the target is applying a model to an unfamiliar experimental situation, factual recall alone is too narrow. The practice should include reconstruction and changed-condition application.

What support is allowed? For a retrieval check, notes may be closed initially, then reopened for feedback. For a real Science reasoning task, the diagram or data table may legitimately remain because interpreting evidence is part of the capability rather than a memory test.

The research therefore informs the practice condition without dictating one universal quiz format. Local evidence must still show whether the learner remembers and uses the Science knowledge better after delay.

17. Research-to-tuition case: interleaving from Mathematics to English

Interleaving research often concerns discrimination among problem types or categories. A Mathematics tutor can see an obvious use: mix several familiar methods so the learner must choose rather than follow a chapter label.

Can the same evidence justify mixing narrative writing, situational writing, grammar correction and comprehension questions randomly in English?

Not automatically.

The mechanism matters. In Mathematics, interleaving may create repeated decisions about which method fits structurally similar problems. In English, mixing tasks can create useful switching practice if the learner genuinely needs to identify different response demands. But random mixing can also fragment extended reading or writing tasks whose value depends on sustained attention and coherent production.

The tutor should transport the discrimination principle, not the superficial schedule. Ask which neighbouring English tasks are commonly confused and whether mixed practice forces a meaningful choice. The local evidence should show better selection or transfer—not merely more variety.

18. Research-to-tuition case: high-impact tutoring and a three-student private class

The research base on tutoring is encouraging, but implementation varies widely.

EEF’s Making a Difference with Effective Tutoring emphasises positive average evidence while warning that not every study is positive and that monitoring and implementation matter. The National Student Support Accelerator publishes quality standards around tutor consistency, training, instruction, data use, relationships and learning integration.

A private three-student class can learn from those principles. It cannot simply inherit their programme effect sizes. Frequency may be lower. Learners may attend by family choice rather than school assignment. Tutor selection may differ. Curriculum alignment may be informal. The group may contain students from different schools. There may be no central coaching system.

The responsible claim is therefore local: small-group tuition can be designed using features supported or recommended in the broader tutoring evidence base, while the actual learner outcomes must be monitored in the local model. “Research-informed” is often the correct phrase where “research-proven for this exact model” would be false.

19. Research-to-tuition case: scaffold fading from literacy to Additional Mathematics

AERO’s current Scaffold Practice guidance describes planned and responsive supports that are gradually removed as student proficiency develops.

The broad principle is plausible in Additional Mathematics: worked examples, partial steps, prompts or representations can support learning and later be reduced.

But the actual scaffold matters. Fading a writing sentence frame and fading an algebraic representation are not identical operations. A representation that carries important mathematical structure may remain useful even for experts. A formula sheet may be an intended resource in some tasks. A diagram can be part of the mathematics, not merely temporary help.

The tutor should therefore ask what cognitive work the scaffold currently performs. Does it make the relationship visible while the learner is building the concept, or does it now select the method on the learner’s behalf? Remove answer-giving assistance when independent performance is the target; retain legitimate tools and representations where they belong to the mature practice.

The research principle travels. The fade schedule and object of fading need subject-specific judgement.

20. Research-to-tuition case: an AI study

Imagine a new study reports that students using an AI assistant produce higher-quality written answers than students without the tool.

Before telling Beatrice to use AI for composition practice, the tutor should reconstruct the study. Did the outcome measure final artifact quality or later unaided writing? Did the students use AI for idea generation, sentence revision, feedback or full drafting? Were prompts taught? How expert were the participants? Was there a delayed transfer test? Did the comparison group have another form of feedback?

If the study only establishes better assisted artifacts, the result may be highly relevant to tool-supported production and weak evidence for independent writing improvement.

The tutor can still use the study to generate a careful hypothesis: perhaps AI feedback can improve revision if the learner remains responsible for diagnosis and rewriting. Then design a bounded local trial and require a fresh unaided paragraph later. The research informs the question; the local receipt determines whether the capability survived the tool.

21. Direct evidence outranks analogy

Education writing loves analogies.

Athletes use deliberate practice, so students should. Pilots use checklists, so learners should. Engineers use redundancy, so studying should. Software teams use version control, so learning plans should.

These analogies can clarify a principle. They do not establish educational effectiveness.

If direct educational evidence exists, use it first. A sports analogy can illustrate why feedback and repeated performance might matter; it cannot substitute for research on feedback in learning. An engineering analogy can explain why a single point of failure is risky; it cannot prove that a particular tutoring routine improves attainment.

The further the evidence comes from education, the more modest the claim should become. Analogy is an idea generator and explanatory tool, not a causal bridge.

22. Practice guides are not the same as primary studies

Strong public guidance often synthesises research, professional knowledge and implementation judgement. That is valuable. It also means the tutor should understand what kind of source they are using.

A practice guide may say an approach is supported by evidence but not reproduce every population boundary on the webpage. A systematic review may combine studies with varied implementations. A single randomised study can have strong internal validity and narrow context. A programme case study can describe implementation richly without establishing causation. A preprint can be current but not yet peer reviewed.

The tutor’s claim should match the evidence class. “This is recommended in current research-informed guidance” is different from “multiple high-quality trials in closely related settings show a consistent effect.”

The National Student Support Accelerator’s explicit categories—research-based, research-informed and emergent—are helpful because they make evidence strength part of the recommendation rather than hiding it behind a single quality label.

23. Context fit should be argued, not felt

“This seems relevant to my students” is a reasonable beginning and a weak ending.

AERO’s 2026 resource Assessing Whether Evidence Is Relevant to Your Context exists because evidence may come from different communities, schools, services or learner groups. Its purpose is to support deliberate reflection about whether an approach is likely to be relevant locally.

The tutor should make the fit explicit. “The study used secondary students with prior exposure to the methods; our learner has comparable prerequisite knowledge. The practice decision—choosing among familiar methods—is the same. Our group size and frequency differ, so I am not importing the effect estimate. I am using the study to justify a mixed-practice trial and measuring local method selection after delay.”

That paragraph is more useful than “research says interleaving works”. It tells the reader which similarities matter and which differences limit the claim.

24. Use a bounded trial when evidence is indirect but promising

Indirect evidence does not force a yes-or-no decision.

If a research-informed practice has a plausible mechanism, low downside and reasonable fit, it can often be tested in a bounded, reversible way. The Trial Run already owns the operational design: baseline, narrow change, predicted receipt, cost signal, stability window and rollback rule.

The Applicability Check determines how confidently the research can justify entering that trial.

A highly direct evidence base may justify moving quickly to a careful implementation. A more distant evidence base may justify only a small pilot. A speculative analogy may justify observation or further reading rather than changing the learner route at all.

Evidence distance changes the size of the first commitment.

25. Define a local receipt before implementation

If the research is supposed to help this learner, what should happen here?

Not “engagement should improve” unless engagement is the target and is defined. Not “learning should be better.”

Use a learner receipt that matches the proposed mechanism. If mixed practice is intended to improve method selection, measure independent selection on fresh mixed tasks. If spacing is intended to support retention, use delayed retrieval. If scaffold fading is intended to strengthen independence, check performance with less answer support. If a feedback routine is intended to improve self-correction, look at the next attempt before feedback returns.

The local receipt does not re-prove the external research. It tells the tutor whether the imported principle is behaving usefully in this learner’s route.

26. Define a falsifier or stop rule

Evidence-informed practice can become ideology when failure is always explained away.

“Retrieval works; she just needs more of it.”

“The scaffold should fade; he is just resistant.”

“Interleaving works; the score drop means desirable difficulty.”

Those explanations may sometimes be true. They cannot be automatic.

Before implementation, state what result would weaken the local hypothesis. If mixed practice repeatedly produces confusion because the component procedures are not yet understood, return to instruction or blocked practice. If prompt fading causes repeated stalled starts without later recovery, the fade may be premature. If an AI feedback routine improves assisted artifacts but unaided writing does not change, limit the learning claim.

A falsifier keeps research evidence connected to learner evidence.

27. Do not import the effect size

A study reports an average improvement. That number belongs to the study unless a defensible synthesis supports broader transport.

Do not tell a parent, “This method improves learning by 20%,” because one paper reported a 20% difference on a particular outcome. Percentage language is especially easy to misread: twenty percent relative to what baseline, on what scale, after what duration, with what uncertainty?

The practical tutoring claim can be useful without importing the number. “This practice is supported by research for improving retrieval under related conditions; we are testing whether it improves your child’s delayed recall and application.”

Mechanism-informed humility is stronger than borrowed precision.

28. Do not use average effects as individual promises

Even when a meta-analysis or large programme shows a positive average effect, individuals vary.

The learner in front of you may benefit more, less or not at all. Implementation may differ. The comparison condition may be stronger or weaker. The learner may already use the strategy. The active bottleneck may sit somewhere else.

Average evidence helps set expectations about a practice class. It does not permit a guarantee.

This is one reason EEF’s tutoring guidance emphasises monitoring and evaluation even though tutoring is positively evidenced on average. Strong average evidence is a reason to take an approach seriously, not a reason to stop observing the learner.

29. The tutor classifications ask different applicability questions

The Tutor Classification Model describes professional functions, and each function reads research through a different decision.

A Homework Helper asks whether research on routines and task management fits the learner’s actual home and school constraints. An Explainer asks whether evidence on examples or representations matches the concept and novice state. A Drill Builder examines dose, spacing, variation and what the practice requires the learner to retrieve or choose. A Diagnostic Tutor asks whether research findings help separate causes rather than merely name symptoms. A Route Designer looks for prerequisite and sequencing implications. A Performance Coach checks whether evidence concerns the same time pressure, support and output conditions. A Learning Architect examines whether an intervention that worked inside one system can coexist with the learner’s school, family and existing supports.

The research paper does not change. The decision question changes what context features become important.

30. Repair, Alignment and Frontier require different evidence transport

Repair mode can often use relatively narrow mechanism evidence. If a learner has a clearly isolated retrieval problem, research on retrieval and spacing may inform the repair even if the study subject differs, provided the tutor still checks local retention and transfer.

Alignment mode requires more local institutional fit. The learner must operate inside current school curriculum, response conventions and assessment conditions. Overseas research can inform learning mechanisms, but current Singapore-specific syllabus or examination claims need local official sources when relevant.

Frontier mode may draw on broader evidence because the goal is extension beyond immediate school requirements. But the tutor still needs a mechanism and a learner benefit. “Advanced” should not mean importing an interesting university activity whose prerequisites, purpose and developmental fit are wrong.

Mode changes how direct the evidence needs to be for the decision at hand.

31. A three-student tutorial can expose applicability differences inside one room

A research-informed routine may fit one learner better than another.

Imagine a mixed-practice routine introduced to three learners. Alicia already knows the component methods and benefits from having to discriminate. Beatrice is still learning one component and becomes confused by rapid switching. Ciara understands the methods but needs more time to explain why each applies.

The same “evidence-based strategy” therefore creates different local jobs. The tutor should not treat programme-level evidence as a group command.

In a three-student class, application may need learner-specific dose and timing even when the shared lesson uses the same broad principle. One learner can enter interleaving while another remains temporarily blocked. One can fade a scaffold while another keeps it. The local route follows readiness, not a slogan.

32. Parent communication: explain evidence without performing certainty

Parents often encounter educational research through confident headlines.

“Testing improves memory.”

“Small-group tuition adds months of progress.”

“Handwriting is better than typing.”

“AI improves writing.”

A tutor should not respond with either dismissal or marketing.

Use a three-part explanation. First: what the strongest relevant evidence actually says. Second: what is different about this learner or setting. Third: how the tutor will test the application locally.

There is good evidence that retrieval can strengthen delayed memory for learned material. Your child’s problem appears to be forgetting concepts she already understands, so the mechanism is relevant. The studies do not guarantee an effect for this exact PSLE Science route, so we are adding short retrieval returns and checking whether recall and changed-question use improve after delay.

That is evidence-informed communication without false certainty.

33. Learner communication: “why are we trying this?”

Older learners can participate in applicability decisions.

“We are trying mixed practice because your methods are accurate when the chapter tells you what to use, but you are still unsure when different methods appear together. Research suggests mixing can help method discrimination in related tasks. We are testing whether that is true for your algebra. If selection gets worse because one method is not yet stable, we will separate it again.”

This explanation gives the learner a mechanism, a reason and a stop condition. The learner is not told to obey “the research”. They are invited to notice what the change does.

That supports metacognition and protects against strategy superstition. A good learning method is not a ritual to perform because adults say it is scientifically proven. It is a tool whose job and boundaries can be understood.

34. The Applicability Check card

  • Claim: What exactly did the research find or recommend?
  • Evidence class: Primary study, synthesis, practice guide, programme evidence, expert consensus, emergent practice or analogy?
  • People: Who was studied and how do they differ from this learner?
  • Knowledge: What subject, task and prerequisite structure was involved?
  • Place: What institutional and cultural setting surrounded the intervention?
  • Provision: What dose, group size, materials, training and support were used?
  • Performance: What did participants actually have to do?
  • Proof: What outcome was measured, under what conditions and after what delay?
  • Pathway: What mechanism is supposed to produce the effect?
  • Context distance: Which differences are likely to matter causally?
  • Claim strength: Research-based expectation, research-informed hypothesis, emergent idea or analogy?
  • Local trial: What smallest implementation can test the mechanism safely?
  • Receipt: What learner evidence should improve if the application fits?
  • Cost signal: What could worsen?
  • Falsifier: What result would make us stop, shrink or redesign the intervention?

Again, this card is not a validated external-validity instrument. It is a tutor decision aid designed to make hidden assumptions visible.

35. Research foundation: relevance and local application

AERO’s evidence-use series is unusually explicit about the transition from evidence quality to local action. Assessing Research Evidence, last updated in June 2026, emphasises both reliability and relevance. Applying Research Evidence, also updated in June 2026, describes deciding whether and how to implement an approach as an ongoing process requiring careful reflection.

AERO also publishes a dedicated 2026 resource on assessing whether evidence is relevant to your context. It explicitly recognises that research may have been generated in different communities, schools, services or learner groups.

The Applicability Check is narrower than these general educator resources. It focuses on one-to-one and small-group tutoring decisions. But the core principle is aligned: rigorous evidence still needs contextual reasoning before it becomes local practice.

36. Research foundation: generalizability is a design problem too

The IES Standards for Excellence in Education Research treat generalizability as something researchers should design for rather than assume afterward. The standards call for intentional sampling or other methods that support inference to populations of interest and for reporting baseline sample characteristics. For development and impact evaluations, they also emphasise implementation and scale.

The IES/NCEE guide Enhancing the Generalizability of Impact Studies in Education goes further into defining target populations, selecting and recruiting samples, assessing sample-target differences and reporting generalizability.

These are research-design standards, not instructions for tutors to run statistical transport analyses. They support a useful humility: generalization is not the automatic reward for a rigorous study. It depends on the relation between the studied sample, the target population and the conditions that make the intervention work.

37. Research foundation: implementation is part of applicability

EEF’s A School’s Guide to Implementation emphasises behaviours, contextual factors and structured processes for embedding evidence-informed approaches. The core warning travels well to tuition: an idea can be sound in principle and fail in day-to-day use if the implementation does not reproduce the important elements or fit the local context.

Implementation is therefore part of the applicability argument. A tutor cannot say a practice “does not work” after using a materially different version, just as the tutor cannot claim the research validates the local version merely because the same label is used.

Name what is core, what is flexible and what local adaptation changed. Then read the learner evidence.

38. Common failure: reject anything not studied locally

This sounds cautious and can become anti-evidence.

Education cannot wait for an exact study of every learner, subject, school and tuition arrangement. Many robust mechanisms and well-supported practices will necessarily be applied outside the exact original sample.

The correct response to context difference is not automatic rejection. It is to identify whether the difference is likely to modify the mechanism or implementation. A study from another country can be highly relevant to memory retrieval. A study from the same country can be poorly relevant if it targets a different capability, age or intervention dose.

Locality is one variable, not a substitute for reasoning.

39. Common failure: copy the intervention surface

A research paper used flashcards, so the tutor uses flashcards.

The important ingredient may have been retrieval plus spaced return plus feedback—not the cardboard rectangle.

A study used peer discussion, so the tutor adds discussion. The active mechanism may have been explaining and comparing reasoning under a structured prompt, not conversation itself.

A programme used a dashboard, so the tutor builds a dashboard. The important element may have been timely use of progress evidence, not the interface.

Applicability improves when the tutor transports the functional mechanism rather than the visible artifact.

40. Common failure: claim the effect before reproducing the conditions

A tutoring programme worked with frequent sessions, trained tutors, aligned curriculum and close progress monitoring.

A tutor copies one feature—small group size—and claims the same evidence base.

That is a category error.

Complex interventions are bundles. Some elements may be essential, some supportive and some incidental. Until the active ingredients are known, removing several elements should weaken the confidence that the original effect will travel.

The local model may still be excellent. It needs its own evidence claims.

41. Common failure: “research says” without source class

A blog summarises a university press release that summarises one study. A tutor says “research says”.

Compressing source classes hides uncertainty.

Prefer the strongest accessible source. Read the study or high-quality synthesis when the decision matters. Use authoritative practice guidance for implementation recommendations. Check whether a claim is current and whether corrections or later evidence have changed it.

If only a practitioner example exists, say so. Emergent practice can still be worth trying when the stakes are low and the mechanism is plausible. It should not wear the language of robust causal evidence.

42. Common failure: confuse “not proven here” with “probably useless”

Absence of direct local evidence is uncertainty, not negative evidence.

If no study has tested a precise three-student Singapore tuition routine, we cannot claim it has no value. We also cannot claim its outcomes are established.

The appropriate response is proportional: use stronger evidence from adjacent contexts to design intelligently, make the mechanism explicit, monitor local receipts, preserve limitations and avoid guarantees.

Evidence discipline is not pessimism. It is the ability to act without pretending uncertainty disappeared.

43. Common failure: make the learner prove the research

The tutor becomes committed to a fashionable intervention.

When the learner struggles, the tutor says the implementation needs more time. When the learner improves, the tutor credits the intervention. Every outcome preserves the original belief.

That is not evidence-informed practice. It is an unfalsifiable story.

The learner is not responsible for making the research true. The tutor is responsible for checking whether the mechanism appears under local conditions and changing course when it does not.

44. The tutor’s Applicability Check

  • What exact claim am I taking from the research?
  • What kind of evidence supports that claim?
  • Who was studied?
  • What subject knowledge and task were involved?
  • What setting surrounded the intervention?
  • What dose and implementation support were provided?
  • Who delivered the intervention and with what preparation?
  • What did participants actually do?
  • What outcome was measured?
  • When was the outcome measured?
  • What mechanism is supposed to connect the intervention to the outcome?
  • Which differences between the study and my learner could change that mechanism?
  • Which differences are probably superficial?
  • Am I importing a mechanism, a practice recommendation or an effect size?
  • What strength of claim is justified?
  • What bounded local trial would preserve safety and interpretability?
  • What learner receipt should improve?
  • What cost signal will I watch?
  • What result would weaken the local hypothesis?
  • What official local source is needed for any Singapore-specific syllabus or assessment claim?

45. The parent’s Applicability Check

  • What does the research actually show?
  • Was it studied with learners like my child?
  • Is the task or subject similar?
  • Was the intervention delivered more often or with more support than our tuition?
  • Did the study measure the same outcome we care about?
  • Is the tutor promising an average effect as an individual result?
  • What part of the research is being applied here?
  • How will we know whether it helps this learner?
  • What would make us stop or change the approach?
  • Is the tutor using research to guide judgement or using “research says” to end the conversation?

46. The learner’s Applicability Check

  • What new learning method are we trying?
  • What problem is it supposed to solve for me?
  • What part of the method comes from research?
  • What part has been adapted for my subject or schedule?
  • What should I notice if it is working?
  • What might feel harder at first without necessarily being bad?
  • What would show that the method is not helping?
  • When will we review it?
  • Am I learning the target skill, or only becoming good at the routine?
  • Can I eventually use the useful part of the strategy without depending on the tutor to run it?

47. What the Applicability Check is not

  • It is not a statistical generalizability analysis.
  • It is not a substitute for reading high-quality research.
  • It is not permission to cherry-pick studies that agree with the tutor.
  • It is not a demand that every study match the learner perfectly.
  • It is not a reason to reject overseas evidence automatically.
  • It is not a reason to import foreign educational policy automatically.
  • It is not a licence to borrow effect sizes as promises.
  • It is not a scoring system for “evidence fit”.
  • It is not a guarantee that a locally plausible application will work.
  • It is a disciplined bridge between external evidence and a specific tutoring decision.

48. The ethical standard

Research authority can be used well or badly.

Used well, it protects tutors from intuition-only teaching, reveals practices that deserve attention, challenges comfortable habits and offers mechanisms that can improve learner routes.

Used badly, it becomes a rhetorical shield. “Science says” replaces explanation. A paper from a distant context becomes a guarantee. An average result becomes a promise to one family. A fashionable intervention continues after local evidence says the learner needs something else.

The ethical responsibility is not to use only perfect evidence. Perfect evidence for one exact learner rarely exists. The responsibility is to know what kind of evidence you have, how far it must travel, which assumptions carry it, and what local evidence will tell you whether the journey was successful.

A tutor should neither worship research nor localise themselves out of it. Read the evidence, respect its design, preserve its boundaries, understand the mechanism, compare the contexts, scale the claim to the distance, and let the learner’s own work decide whether the imported idea earns a place in the route.

Evidence and connected reading

Final compression

Start with the real research claim.

Who was studied? What did they do? Where? How often? With what support? Who delivered it? What was measured? When?

Then look at the learner in front of you.

Which differences could change the mechanism? Which are probably superficial? Is the subject structure similar? Is the local dose faithful enough? Does the tutor have the professional skill the intervention assumes? Does the outcome match the outcome you actually need?

Transport the mechanism carefully. Do not borrow the effect size. Downshift the claim as context distance grows. Use a bounded trial. Define the receipt. Define the stop rule. Let local evidence update the route.

The Applicability Check turns “research says” into the question a tutor actually needs: what does this evidence justify us trying, expecting and claiming for this learner, under these conditions, now?

That is the Applicability Check.

That is Tutor Handbook Volume 0071.