The Tutor Handbook · Volume 0106 · Series ID THB-0106
The Tutor Handbook: Complete series index.
A tutor sees a teaching idea online at 11.40 p.m.
The post is confident. It has a chart. The comments are enthusiastic. A school leader says it transformed revision. A vendor page says the method is “research-backed”. A short video shows smiling learners. An AI summary says the evidence is strong.
The tutor has a learner tomorrow who is struggling with exactly the problem the post claims to solve.
Should the tutor try it?
The Evidence Source Gate is the decision before adoption or trial in which a tutor asks what kind of source is making the claim, what evidence actually sits behind it, how much confidence that evidence deserves, what the source does not establish, and whether the idea is important enough to justify further checking.
This volume is deliberately upstream of two existing Tutor Handbook jobs. The Applicability Check asks whether evidence from another age, subject, country or setting belongs in this learner’s route. The Trial Run asks how a new route should be tested before becoming standard practice.
The Evidence Source Gate asks the question that comes first: is there enough trustworthy substance here to spend professional attention on the idea at all?
Quick Answer
Do not treat all educational claims as equal because they arrive in the same browser window. A systematic review, a single experiment, an official practice guide, a vendor case study, a tutor’s anecdote, a social-media post and an AI-generated summary can all be useful, but they carry different evidential weight and different failure modes.
Before changing tutoring practice, identify the source class, trace important claims to the original evidence where possible, check currency, authority, accuracy, purpose and conflicts of interest, separate direct education evidence from analogy, look for contrary findings and limits, and decide how much confidence is justified. Low-confidence evidence may still generate a small reversible hypothesis. It should not be promoted into a universal rule.
1. “Research-Backed” Is Not a Research Design
Educational marketing often compresses a long evidence chain into two words: research-backed.
The phrase can mean several very different things. The product may have been tested in a randomised trial. Its individual features may resemble practices supported by research. A company may have surveyed satisfied users. A white paper may cite studies of neighbouring interventions. A founder may have a plausible theory. The claim may be accurate, stretched or impossible to evaluate from the page.
The tutor’s first move is therefore translation. Replace the label with a concrete question: what evidence supports which exact claim?
If the claim is “students like the tool”, a satisfaction survey may be relevant. If the claim is “the tool improves mathematics achievement”, satisfaction is insufficient. If the claim is “this explanation reduces misconceptions”, the evidence should involve misconception-related outcomes, not only time on task or completion.
The source gate begins by making the claimed outcome explicit.
2. Source Class Matters Because Failure Modes Differ
A useful tutor does not need to become a research methodologist before every lesson. But recognising source classes prevents basic category errors.
- Systematic reviews and meta-analyses can synthesise multiple studies, but their conclusions depend on the quality and comparability of the included evidence.
- Individual experiments can provide stronger causal evidence for a specific intervention and setting, but one study may not generalise widely.
- Observational studies can identify patterns and associations without automatically proving cause.
- Official practice guides often synthesise evidence into usable recommendations, but recommendations can combine research, policy and professional judgement.
- Vendor studies and case studies can provide useful implementation detail, but incentives and selective reporting deserve attention.
- Teacher or tutor anecdotes can surface practical hypotheses and local context, but they cannot establish general efficacy.
- Social-media posts and videos can reveal ideas quickly, but brevity often removes methods, comparison groups, unsuccessful cases and uncertainty.
- AI summaries can help navigate material, but important claims should be checked against the source rather than treating the summary as the authority.
The point is not to rank every source on one ladder. A qualitative case study may answer an implementation question better than a meta-analysis. The point is to match the source to the claim.
3. Begin With the Original Claim, Not the Screenshot of the Claim
Educational ideas travel through layers. A journal article is summarised by a university news page. That summary appears in a newsletter. A consultant turns it into a slide. A social-media post screenshots the slide. An AI system summarises the social-media discussion.
By the time the tutor sees the claim, qualifiers may have disappeared.
If the claim matters to a teaching decision, move laterally and backward. Find the original paper, official guide or primary source. Check whether the source really says what the derivative post says. Look at the population, outcome, comparison and limitations.
This does not mean every small idea requires an hour of research. Verification depth should match consequence. A low-risk warm-up variation can be tried cautiously with limited evidence. A major route change, expensive product purchase, high-stakes assessment interpretation or claim about a learner’s capability deserves stronger sourcing.
4. Currency: Is the Claim Still About the Current World?
Some educational mechanisms are durable. Other facts change quickly.
A classic study on memory may remain conceptually useful. A claim about the current PSLE format, a software feature, an AI model’s capabilities, a privacy policy or an examination regulation needs recent verification.
AERO’s resources for evaluating non-academic sources include Currency as one part of the CRAAP framework. That is especially useful for tutors because many practical claims arrive through webpages rather than journals.
Currency is not “newer is always better”. A new blog post can still rely on weak evidence. An older high-quality study can remain important. The question is whether time changes the truth conditions of the claim.
5. Relevance: Does the Evidence Address the Actual Tutor Question?
A source can be excellent and still answer the wrong question.
A study showing that retrieval practice improves retention does not automatically tell a tutor how often a particular learner should use flashcards. Research on whole-class instruction may not establish the best routine inside a three-student tuition room. A university study with adult participants may illuminate a mechanism without directly validating a Primary 4 protocol.
The source gate therefore asks whether the measured outcome is the outcome the tutor cares about. Did the study test immediate performance or delayed retention? Accuracy or speed? Learner confidence or objective achievement? Completion or independent transfer?
Once the source itself survives, The Applicability Check handles the deeper transfer question. Do not collapse source credibility and contextual applicability into one judgement.
6. Authority: Who Is Making the Claim and What Can They Actually Know?
Authority is not the same as fame.
An education ministry is authoritative for its own current examination rules. It is not automatically the strongest source for a general cognitive-science claim. A university researcher may be highly qualified in memory but not in Singapore curriculum policy. A vendor may know its software architecture better than an independent evaluator while having a commercial incentive when describing effectiveness.
Ask whether the source has direct access to the relevant evidence and appropriate expertise for the claim. Also distinguish institutional authority from evidential authority. An official body can publish a policy decision that is authoritative as policy even when the policy’s effectiveness remains an empirical question.
7. Accuracy: Can the Claim Be Checked Against Evidence?
Good educational sources usually allow a reader to inspect where claims came from. They cite studies, describe methods, name data sources or state the basis for recommendations.
Warning signs include precise effect claims without sources, graphs without axes or sample descriptions, universal statements based on testimonials, and summaries that quote research language without identifying the research.
Accuracy checking also includes internal coherence. Does the source claim a large improvement while showing only a satisfaction survey? Does the evidence compare the intervention with no intervention when the practical choice is between two active methods? Does a graph use percentages without denominators? Is an “average gain” being presented as though every learner improved?
The tutor does not need to solve every statistical problem. Enough inspection can often reveal whether the source supports the strength of language being used.
8. Purpose: What Is the Source Trying to Make You Do?
Purpose shapes presentation.
A vendor wants a purchase. A policy body wants implementation. An advocacy group wants attention to a problem. A researcher wants to answer a question and publish findings. A social-media creator may want reach. A tutor sharing a successful lesson may simply want to help colleagues.
None of these purposes automatically invalidates a source. Commercial evidence can be accurate. Advocacy can be evidence-rich. Academic work can contain bias and error. Purpose matters because it suggests which missing information deserves extra checking.
If a source benefits when the reader adopts an intervention, look for independent evidence. If a post presents only success stories, ask what happened in unsuccessful cases. If a policy guide recommends a practice, distinguish the recommendation from the underlying empirical strength.
9. Separate Direct Evidence From Analogy
Tutors often learn from other domains. Sport can illustrate deliberate practice. Medicine can illustrate differential diagnosis. Aviation can illustrate checklists. Engineering can illustrate failure modes.
Analogies can sharpen thinking. They do not automatically establish educational effectiveness.
If a tutoring protocol is justified mainly because “elite athletes do this”, the tutor still needs direct education or learning evidence before presenting the protocol as evidence-based. The analogy may generate a hypothesis: perhaps spaced rehearsal matters. The evidence must then come from learning research relevant to that mechanism.
This distinction protects the Tutor Handbook from clever metaphors becoming hidden proof.
10. Vendor Evidence Needs a Different Reading Posture
A tutoring platform says its adaptive system increases achievement. The company provides a case study showing schools improved after adoption.
The tutor should ask several questions. Was there a comparison group? Who selected the participating schools? Were learners already improving? Was the outcome an independent assessment or the platform’s own score? How many learners stopped using the tool? Were results published for all sites or a selected subset? Was the analysis conducted independently?
A positive answer to these questions increases confidence. Missing answers do not prove the tool is ineffective; they limit what the evidence can establish.
When the cost and risk are low, limited evidence may still justify a reversible trial. The tutor should describe it honestly: “promising enough to test here”, not “proven to work”.
11. Social-Media Evidence Is Often a Discovery Layer
Short-form platforms can be excellent at surfacing practical ideas. A thirty-second clip can reveal a questioning routine the tutor has never seen. A teacher thread can point to a useful paper. A community discussion can expose implementation problems that formal publications barely mention.
Use these sources for discovery, not automatic authority.
The more consequential the claim, the further the tutor should move from the post toward the evidence. “This worked well in my class” is valuable testimony about one context. “This method improves memory for everyone” is a general claim requiring more.
Popularity is not replication. A teaching idea can travel because it is intuitive, visually attractive or easy to explain rather than because it has survived careful comparison.
12. AI Summaries Are Maps, Not Final Citations
AI systems can rapidly summarise a research area, identify terminology and suggest sources. That is useful at the search stage.
But a tutor should inspect material claims in the source itself. Summaries can omit caveats, merge findings from unlike studies, misstate publication details or present an ongoing pilot as established evidence.
The AI Material Verification Gate applies here too. AI can help the tutor reach the source more quickly. It should not become the source when the original evidence is available.
13. A Worked Composite Case: The “Brain-Based” Revision Method
This is a constructed case.
A parent sends a video claiming that rewriting notes in a particular colour sequence “activates both sides of the brain” and doubles retention. The presenter uses scientific vocabulary and displays a brain image.
The tutor does not mock the parent or accept the claim. The source gate begins.
What is the claim? A specific colour routine improves retention. What evidence is cited? None beyond testimonials. What mechanism is asserted? A simplistic hemispheric explanation. What outcome was measured? Unclear. Is there direct learning evidence? Not supplied.
The tutor may still notice a useful adjacent idea: organised notes can support retrieval if learners later practise recalling and using the content. But that is a different claim supported by different evidence.
The source gate prevents a weak neuroscience story from smuggling a whole study routine into the learner’s route.
14. A Worked Composite Case: A Strong Study With the Wrong Outcome
A tutor reads a well-designed experiment showing that a particular digital practice system increases task completion. The sample is large, the methods are transparent and the result is statistically convincing.
The tutor’s problem is different. One learner already completes everything but fails to transfer the method to unfamiliar questions.
The source survives quality checks and still does not answer the tutor’s question. High-quality evidence about the wrong outcome remains the wrong evidence for this decision.
This case shows why the gate asks both reliability and relevance before moving to applicability.
15. A Worked Composite Case: The Vendor With Independent Evidence
A vocabulary platform makes a commercial claim and links to an independently conducted study. The paper describes the sample, comparison condition, outcome measure and limitations. Effects appear positive for learners similar in age to the tutor’s students.
The fact that the company benefits commercially does not erase the evidence. It tells the tutor to inspect independence, outcomes and replication carefully.
The source gate may conclude: moderate confidence that the tool can help the measured vocabulary outcome under similar conditions. Then the Applicability Check asks whether those conditions match the present learner. If promising, The Trial Run can define a bounded test.
The source is neither accepted nor rejected because of who sold the product. It is evaluated because of what the evidence can support.
16. Confidence Should Be Graduated, Not Binary
Educational decision-making becomes brittle when evidence is labelled simply “proven” or “not proven”.
AERO’s evidence decision-making resources use levels of confidence and explicitly treat professional judgement as necessary. That is a useful stance for tutoring. A hypothesis with plausible mechanism and anecdotes may deserve low confidence. Several rigorous studies with relevant outcomes may justify higher confidence. A practice guide synthesising a mature evidence base can support stronger initial expectations.
Confidence should also attach to the claim, not the brand. A program may have strong evidence for one outcome and weak evidence for another. A paper may provide strong causal evidence in one population while leaving transfer uncertain.
This allows tutors to act without pretending certainty.
17. The Evidence Source Card
- Claim: What exactly is being claimed?
- Source class: Review, experiment, observational study, practice guide, case study, vendor material, anecdote, social post or AI summary?
- Original source: Can I trace the claim to the underlying evidence?
- Currency: Does time materially affect the claim?
- Authority: Does the source have relevant expertise and direct access to the information?
- Accuracy: Are methods, data or references available for checking?
- Purpose: What does the source want the reader to believe or do?
- Outcome match: Did the evidence measure the result I care about?
- Directness: Is this education evidence or an analogy from another field?
- Contrary evidence: What would a sceptical search find?
- Limit: What does the source explicitly not establish?
- Confidence: What strength of action is justified: ignore, investigate, small trial, or stronger adoption consideration?
The card is not a scoring formula. Different questions make different fields more important. An official current policy page may outrank a five-year-old study for an examination rule. A systematic review may outrank a vendor testimonial for an efficacy claim.
18. Do Not Average Unlike Evidence Into a Fake Score
A tutor finds one systematic review, three enthusiastic blog posts and six parent testimonials. It would be meaningless to say the idea has “ten pieces of evidence”.
Evidence differs in independence, design, outcome and relevance. Count quality and fit conceptually rather than adding sources as though each were one vote.
This mirrors The Evidence Triangulation Check. Multiple evidence streams can strengthen understanding when their jobs are kept distinct. They become misleading when collapsed into one invented number.
19. Strong Evidence Does Not Remove the Need for Local Observation
Even when a practice has strong research support, implementation can fail. The tutor may use the wrong dose, apply the method to the wrong learner problem, remove a necessary support, or measure the wrong outcome.
Research evidence gives a better prior expectation. It does not eliminate the learner.
This is why the Tutor Handbook separates source, applicability, trial, implementation fidelity and later evidence. A strong source can justify trying an idea. The learner’s performance still decides whether the route is working here.
20. Weak Evidence Does Not Always Mean Do Nothing
Some practical questions have thin research. Tutors still have to teach tomorrow.
When evidence is weak but the proposed action is low-risk, reversible and educationally plausible, a small trial can be reasonable. The tutor should lower the claim, not necessarily abandon action.
For example: “There is not strong evidence that this exact note-formatting routine improves transfer, but it may reduce organisational friction for this learner. We can try it for two weeks while keeping the actual retrieval and transfer checks unchanged.”
Low confidence demands stronger humility and clearer stop conditions.
21. High-Risk Claims Need Stronger Sources
The evidential threshold should rise with consequence.
A new warm-up structure can be tested cheaply. A claim that a learner has a medical or psychological condition cannot be supported by ordinary tutoring observations. A major curriculum change, expensive long-term program, removal of legitimate accommodations or strong public claim about effectiveness needs much stronger authority and evidence.
The source gate therefore includes safety and scope. Good tutors know when the question should leave tuition rather than search harder for confirmation of something they are not qualified to determine.
22. How Current Evidence Organisations Frame the Problem
The Australian Education Research Organisation’s guide to looking for research evidence explicitly notes that some sources are more credible than others and that recognising credible sources increases the chance of finding high-quality evidence.
AERO’s Assessing Research Evidence guide focuses on reliability and relevance, and its CRAAP resource uses Currency, Relevance, Authority, Accuracy and Purpose to structure judgement about webpages and other non-academic sources. AERO explicitly says the tool does not replace professional judgement.
Its Applying Research Evidence guide treats implementation as an ongoing process after rigorous and relevant evidence has been identified. That sequence closely matches this volume’s boundary: source quality first, contextual applicability next, then implementation.
The National Student Support Accelerator’s Tutoring Quality Standards make another useful distinction by labelling standards as research-based, research-informed or emergent. The categories prevent a practitioner consensus from being described as though it came from the same evidence base as a robust research finding.
23. A Tutor’s Five-Minute Source Check
When time is limited, a tutor can still avoid the largest errors.
- Write the claim in one sentence.
- Open the original source rather than relying on the summary.
- Check who and what was studied and what outcome was measured.
- Look for limitations or an independent source that disagrees.
- Decide what confidence level the evidence earns and match the size of action to that confidence.
Five minutes will not resolve a complex literature. It is often enough to stop a headline from becoming a teaching rule.
24. The Longer Check for Major Changes
For consequential changes, slow down.
Search for reviews. Read methods and limitations. Check whether results replicate. Look for independent evaluation. Inspect whether the measured population resembles the learner. Check implementation demands. Compare the proposed intervention with what the learner is already receiving. Ask what would be displaced. Define a trial and a failure condition before adopting.
The time spent here is not academic decoration. Major educational changes can consume months of learner opportunity. The stronger the commitment, the stronger the source discipline should be.
25. Common Source-Gate Failures
- Authority by design: a polished PDF is mistaken for strong evidence.
- Popularity by truth: many shares are treated as replication.
- Vendor rejection: commercial evidence is dismissed automatically instead of inspected.
- Vendor trust: commercial evidence is accepted automatically because the product looks professional.
- Summary substitution: a blog or AI summary replaces the original source.
- Outcome mismatch: evidence for engagement is used to claim learning.
- Analogy laundering: research from sport, medicine or business is presented as direct education evidence.
- One-study absolutism: a single positive finding becomes a universal rule.
- Old-current confusion: an outdated page is used for a changing policy or examination fact.
- Evidence paralysis: the tutor refuses low-risk reversible action because perfect research does not exist.
26. What Parents Should Hear
Parents often encounter educational claims before tutors do. They may send an article, product, influencer video or method recommended by another family.
A professional response does not need to be defensive. “This is interesting. I want to see what evidence supports the specific claim and whether it fits your child’s current job.” That keeps the parent inside the reasoning rather than treating outside ideas as interference.
If evidence is weak, explain the uncertainty. If the idea is low-risk and plausible, a bounded trial may be possible. If a strong existing route is already working, the burden for disruption should be higher.
27. The Sequence: Source → Applicability → Trial → Receipt
A clean professional-learning sequence is:
- Source Gate: what is the evidence and how much confidence does it deserve?
- Applicability Check: does the evidence plausibly travel to this learner and context?
- Trial Run: how can the idea be tested reversibly without replacing the entire route?
- Implementation Fidelity: was the approach actually used as intended?
- Learning Receipt: did later learner evidence justify keeping, changing or abandoning it?
Skipping the source gate makes weak claims look ready for implementation. Skipping applicability turns strong research into context-blind prescription. Skipping trial and receipt turns professional learning into belief.
28. The Human Standard
Good tutoring does not require knowing every paper. It requires knowing the difference between a claim and the evidence for the claim.
A tutor should be curious enough to notice new ideas, sceptical enough to inspect them, humble enough to preserve uncertainty, practical enough to run small trials when evidence is incomplete, and disciplined enough to abandon an attractive method when later evidence does not support it.
The question is not whether an idea came from research, a colleague, a company, a parent, a video or an AI system. The question is what kind of evidence survived after the tutor looked past the packaging.
That is the Evidence Source Gate.
That is Tutor Handbook Volume 0106.
Connected Reading and Sources
- The Tutor Handbook | Complete Series Index
- The Tutor Handbook Vol No.0071 | The Applicability Check
- The Tutor Handbook Vol No.0052 | The Trial Run
- The Tutor Handbook Vol No.0079 | The Implementation Fidelity Check
- AERO | The Value of Research Evidence
- AERO | Looking for Research Evidence
- AERO | Assessing Research Evidence
- AERO | Evaluating Non-Academic Sources
- AERO | Applying Research Evidence
- AERO | Evidence Decision-Making Tool
- National Student Support Accelerator | Tutoring Quality Standards