The Tutor Handbook · Volume 0166 · Series ID THB-0166
Series route: The Tutor Handbook — Complete Series Index.
A learner has access to an AI tutor. The account exists. The platform works. The learner logs in. A tutor sits nearby, reminds the learner to start, helps interpret the first prompt and keeps the session moving. Usage rises.
Has learning improved?
Not yet as a conclusion. Access, use, engagement and learning belong to the same pathway, but they are not the same event.
The Human-Supported AI Engagement Gate is the tutor’s discipline of deciding what human support around learner-facing AI is supposed to accomplish, whether that support produces meaningful engagement with the target learning job, and what independent evidence will show that increased use became learner capability rather than merely more screen time.
This is an increasingly important tutoring problem because AI can make support abundant while keeping the educational mechanism ambiguous. A learner can have sophisticated conversational help and still avoid the target operation. A human tutor can increase platform usage and still have no evidence that the learner learned more. Or the human can provide exactly the structure needed for the AI tool to become useful practice.
Quick Answer
Do not judge learner-facing AI by access, enthusiasm or usage alone. Define the learning job before the session. Decide what the human is supporting: starting, persistence, interpreting instructions, choosing an appropriate task, checking an answer, or reflecting on an error. Keep the AI from performing the target operation the learner is meant to own. Then use a fresh response outside the immediate AI interaction to check whether the capability survives.
Human support can be valuable even when its first effect is only engagement. If a learner never uses a potentially useful tool, there is no learning opportunity to evaluate. But engagement is an intermediate outcome. The tutor should not let “minutes increased” quietly become “learning improved”.
The central chain is: access → actual use → meaningful task engagement → appropriate cognitive work → feedback and correction → later performance. Human support may strengthen one or more links. The tutor should know which link they are trying to repair.
1. What This Volume Owns
This volume owns human support around a learner-facing AI tutor or AI-enabled practice tool.
It does not replace the AI Material Verification Gate, which owns the accuracy, validity, curriculum fit and safety of AI-generated questions, explanations and feedback. It does not replace the Support Provenance Check, which asks who or what helped produce learner work. It does not replace the Dosage Differential, which asks whether more tutoring time is needed or the current tutoring is simply the wrong kind. It also does not own AI that supports the human tutor behind the scenes rather than interacting directly with the learner.
The exact decision here is narrower: when a human tutor helps a learner use AI, how do we know whether the human is creating a genuine learning opportunity or merely increasing technology activity?
2. “AI Tutoring” Is Not One Intervention
One reason this problem becomes confused is that very different systems receive the same label.
A learner-facing conversational tutor that asks questions and explains errors is one model. An adaptive practice platform with AI-generated hints is another. A general-purpose chatbot used for homework help is another. A human tutor using an AI copilot that privately suggests questions or prompts is another. A conventional practice platform with an AI wrapper is yet another.
The National Student Support Accelerator’s 2026 synthesis makes this distinction central: AI tutoring is not a monolith, and results from one delivery model should not be casually transferred to another.
Source: National Student Support Accelerator — AI Tutoring Is Not a Monolith: What We Actually Know.
Before discussing whether “AI works”, the tutor should identify what the learner actually encounters and what human support changes in that model.
3. Access Is the First Gate, Not the Final Outcome
A platform can be available without being used. A learner may forget the login, avoid starting, find the interface confusing, feel uncertain about what to ask, or simply choose a faster source of help. Schools and families can therefore spend money on access that produces little meaningful exposure.
Human support can solve real access-to-use problems: setting a routine, clarifying the purpose of the tool, helping the learner begin, choosing the right unit, resolving a technical obstacle, or making the first task less intimidating.
These are legitimate educational implementation jobs. But the receipt is initially only that the learner used the tool under improved conditions. The tutor should resist jumping directly to a learning claim.
4. Use Is Not the Same as Engagement
A learner can remain logged in while doing little. They can click rapidly through hints, ask the AI for direct answers, skim long explanations, or copy a generated solution into school work. Minutes and messages can rise while the target learning operation remains absent.
Meaningful engagement therefore needs a task-level definition. Did the learner attempt before requesting help? Did they answer the AI’s question or simply ask for the solution? Did they compare their reasoning with feedback? Did they revise an error? Did they generate an explanation? Did they complete a fresh problem after support?
The tutor should prefer observable learning actions over platform activity as the immediate engagement measure.
5. Engagement Is Still Not Learning
A learner can engage deeply and still fail to retain or transfer. They may understand an explanation while it is visible but not retrieve the idea later. They may complete the AI dialogue successfully because the tool supplied cues that disappear elsewhere.
This is not unique to AI. A learner can be highly engaged in a tutor-led lesson and still lack independent performance later. AI simply makes the distinction easier to forget because digital systems produce abundant usage data.
Learning claims require later evidence: a fresh question, a delayed return, a changed representation, school-generated work, or another performance where the target operation belongs to the learner rather than the system.
6. A 2026 Working Paper Shows Why Human Support Deserves Separate Measurement
A 2026 EdWorkingPaper by researchers studying AI tutoring in elementary reading reported two randomised experiments with 355 kindergarten to Grade 5 pupils. Human tutors were used to support engagement with the AI system rather than to provide the main academic instruction. The intervention increased AI usage by roughly one to four minutes per week and raised measures of engagement, but usage remained low and the study did not find improvements in reading achievement.
Source: Access Is Not Enough: Human Support Improves Engagement with AI Tutoring, 2026 working paper; see also the National Student Support Accelerator study summary.
This is emergent evidence, not a settled peer-reviewed conclusion, and the population and reading context are specific. Its value for tutors is conceptual: an intervention can genuinely improve engagement while leaving achievement unchanged. Engagement therefore deserves to be measured as an intermediate mechanism rather than treated as a synonym for learning.
7. Another 2026 Working Paper Shows Why the AI Layer Must Be Isolated
A separate two-year cluster-randomised experiment in 18 Tennessee middle schools examined Khan Academy with Khanmigo. The study reported small Mathematics gains from assignment to the AI-enabled programme, but gains were similar in magnitude to those associated with ordinary Khan Academy practice, while direct engagement with the AI tutor itself was relatively infrequent.
Source: One Click Away: AI Tutoring with Khanmigo in a Two-Year School Experiment, 2026 working paper.
Again, this is a working paper and should not be generalised casually. The methodological lesson is valuable: when a platform contains ordinary practice plus an AI layer, the tutor should not attribute the whole platform effect to the AI conversation. Identify which component the learner actually used.
8. Define the Target Operation Before Opening the Tool
“Use the AI tutor for twenty minutes” is an activity instruction, not a learning target.
A stronger session begins with the operation the learner should perform: identify the percentage base, distinguish inference from evidence, explain a causal chain, choose between two algebraic methods, revise a paragraph for coherence, or retrieve a set of concepts without notes.
Once the target is named, the tutor can inspect whether the AI supports the learner in performing it or quietly performs it on the learner’s behalf.
The same AI feature can be helpful for one target and harmful for another. A hint that names the method can support execution practice but destroy evidence of method selection. An AI summary can support checking but replace summarisation practice. A generated example can clarify a concept but remove the learner’s need to create one.
9. Ask Who Performed the Target Cognitive Work
This question is the core of responsible AI-supported tutoring.
If the target is explanation, did the learner explain before the AI evaluated the explanation, or did the AI write the explanation first? If the target is evidence selection, did the learner choose the line, or did the AI point to it? If the target is planning, did the learner prioritise tasks, or did the AI produce the plan?
AI can legitimately scaffold a target operation. Early learners may need prompts, examples or partial structures. The tutor simply needs to know what support remains so later independence can be checked fairly.
10. Human Support Has Different Jobs
“Human support” is itself too broad. A tutor can support AI use in several ways, and each creates a different educational effect.
- Logistical support: login, device setup, locating the correct module, scheduling use.
- Motivational support: helping the learner begin, persist and return after difficulty.
- Interpretive support: clarifying what the AI is asking or what its feedback means.
- Strategic support: helping the learner choose when to ask for a hint, when to attempt again or when to change task.
- Instructional support: teaching missing knowledge the AI interaction has exposed.
- Verification support: checking whether an AI explanation, answer or question is accurate and appropriate.
A programme can say “human-supported AI” while providing mainly logistical support. Another may provide substantial teaching. Those models should not be treated as educationally equivalent.
11. Human Support Should Solve a Named Bottleneck
The tutor should know what problem their presence solves.
If the learner does not start because the platform is confusing, logistical help may be enough. If the learner quits when the AI challenges an answer, motivational support may matter. If the learner cannot tell whether a long AI explanation is relevant, interpretive support may be necessary. If the AI exposes a missing prerequisite, direct teaching may be more efficient than continuing the conversation.
Without a named bottleneck, the human can become an expensive companion to technology rather than an instructional resource.
12. Constructed Case: Alicia Logs In but Does Not Begin
This is a fictional teaching case. Alicia has access to an AI Mathematics tutor but rarely uses it. Her tutor initially assumes motivation is weak. Observation shows a simpler problem: the platform opens to a general chat interface, and Alicia does not know what kind of question to type.
The tutor creates a narrow entry routine: open the current topic, attempt one school question first, then ask the AI only about the exact step that became uncertain. Usage increases because the starting state is clear.
The learning receipt is not the increased message count. After several sessions, Alicia receives a fresh school-style question without AI. She identifies the first valid step independently. That later response tells the tutor whether the structured AI use contributed to a capability that survives the tool.
13. Constructed Case: Beatrice Uses AI to Avoid Evidence Selection
Beatrice uses an AI tutor for comprehension. She asks, “What evidence proves this inference?” The AI points to a strong line and explains why. Beatrice finds the tool helpful and completes assignments quickly.
The tutor notices that the AI is performing the target operation. The session is active, but the wrong party is selecting evidence.
The tutor changes the prompt routine. Beatrice must first choose one line and explain why she thinks it is relevant. Only then can she ask the AI to critique the choice or suggest what would make the evidence stronger.
Now the AI functions as feedback rather than answer selection. A fresh passage later checks whether Beatrice can select direct evidence without the AI.
14. Constructed Case: Ciara Gets Long Explanations but No Better Answers
Ciara asks an AI tutor to explain Science questions. The explanations are detailed and usually accurate. She reads them carefully. Her own written explanations remain descriptive rather than causal.
The tutor concludes that exposure to explanation is not the missing operation. Ciara needs to generate causal links.
The human-supported routine changes: before reading the AI explanation, Ciara writes a three-step causal chain. She asks the AI to identify which link is unsupported. She revises only that link. The tutor then gives a fresh question with the AI closed.
The technology becomes a critic of learner-generated reasoning rather than a replacement for it.
15. Constructed Case: Denise Chases Hints
Denise enjoys an adaptive Mathematics platform. When a question is difficult, she immediately opens the first hint, then the second, then the worked solution. Her usage data is excellent. Her independent method selection is not improving.
The tutor introduces a help threshold. Denise must record one attempted representation or first step before opening a hint. If the attempt is genuinely blocked, she uses the smallest hint available, closes it, and resumes the problem.
The human is not preventing help. The human is protecting a short productive attempt before help substitutes for the decision.
16. Constructed Case: Emily’s AI Study Plan
Emily asks an AI assistant to create her weekly revision plan. The plan is beautifully structured. She follows it for two days, then school priorities change and she does not know how to adapt it.
The tutor identifies that the artefact is stronger than the learner’s planning capability.
Instead of banning AI, the tutor changes the division of labour. Emily chooses the priorities, estimates available time and identifies non-negotiable deadlines. The AI is allowed to check for overload and suggest alternative sequencing. Emily makes the final trade-off.
A week later, she plans a changed schedule without AI first, then uses the tool as a verifier. Human support has shifted the learner from outsourcing planning to using AI as a secondary check.
17. The Engagement Ladder
A practical tutor can think in levels without pretending they form a validated measurement scale.
- Available: the learner has access to the tool.
- Entered: the learner opens the relevant task or interaction.
- Active: the learner responds, attempts, asks or revises rather than merely viewing.
- Targeted: the activity requires the intended learning operation.
- Responsive: the learner uses feedback to change the attempt.
- Independent return: the target capability appears later without the AI performing it.
Moving up the ladder requires different evidence. Login data can show entry. Interaction logs can show activity. Only task content can show whether the target was engaged. Later performance is needed for a stronger learning claim.
18. Measure Meaningful Use, Not Just Time
Time is easy to count and easy to misread.
Five minutes of deliberate attempt, feedback and revision can be more educationally valuable than twenty minutes of passive explanation. Conversely, some complex tasks genuinely require time. The point is not that short sessions are better. It is that time needs a task interpretation.
A tutor can record simple qualitative indicators: attempted before hint, generated explanation, revised after feedback, completed a fresh item, or stopped because prerequisite teaching was needed.
These are not proprietary scores. They are observable events that explain what the time contained.
19. Do Not Reward Message Volume
Conversational AI can make message count look like engagement. A learner may send many short requests because the AI keeps requiring clarification. Another may produce one careful solution and ask one high-quality question.
Message volume is therefore a poor stand-alone learning indicator. The tutor should inspect the function of the messages. Are they attempts, explanations, requests for verification, requests for direct answers, or repeated rephrasing of the same confusion?
Quality of interaction is closer to the educational mechanism than quantity of interaction.
20. Human Support Can Accidentally Create Dependence on the Human
A learner may use AI only when a tutor sits beside them. The tutor reminds them what to ask, interprets every response and decides when to continue. Usage improves, but the implementation has created a new dependency.
If the long-term goal includes learner-managed technology use, the human role should have a fade path. The learner can progressively choose the task, decide when a hint is justified, judge whether an answer is plausible and close the tool when it stops helping.
The tutor should not disappear abruptly. They should transfer specific decisions.
21. Human Support Can Also Protect the Learner From the Tool
Human support is not only about increasing use. Sometimes the correct decision is to stop or redirect an AI interaction.
The tool may produce an inaccurate explanation, introduce content outside the learner’s level, overcomplicate a simple question, reveal a complete solution before the learner has attempted, or continue a conversation when direct teaching would be more efficient.
A tutor can intervene by verifying claims, narrowing the prompt, moving back to a textbook or school source, or closing the AI and teaching the prerequisite directly.
Good implementation is not measured by maximum AI use. It is measured by appropriate use.
22. Protect the Boundary Between Feedback and Answer Production
AI is especially powerful at producing polished outputs. This makes provenance important.
If the learner’s target is writing an explanation, the AI should not write the final explanation before the learner has produced one. If the target is checking an explanation, AI-generated comparison can be useful. If the target is learning vocabulary, an AI example sentence can help, but the learner still needs to interpret and use the word in a fresh context.
The tutor should define where AI input enters the process so later work is not mistaken for independent production.
23. Build a Fresh Return Outside the AI Conversation
The cleanest everyday check is often a short fresh task after the AI session.
Close the conversation. Keep legitimate access supports that are unrelated to the target. Give one new item requiring the same capability. Ask the learner to solve, explain, select or revise without the AI performing the operation.
The return does not need to be difficult. Its purpose is to answer a simple question: did the learner carry anything useful out of the AI interaction?
24. Delayed Returns Matter When Immediate Memory Is Strong
An immediate fresh item can still be influenced by the exact AI explanation that remains active in memory. When the learning claim matters, a delayed check adds information.
The next session, a later school task or a changed representation can show whether the capability remains available when the AI dialogue is no longer vivid.
This is especially important when the AI uses memorable phrasing or provides a complete worked path. Recall of the explanation is not identical to control of the underlying method.
25. Use the AI Interaction as Diagnostic Evidence Carefully
AI conversations can reveal useful things: what questions the learner asks, where they request help, which errors recur, whether they revise after feedback and how long they persist.
But these traces occur inside a support-rich environment. The AI may steer the conversation, phrase the options or rescue dead ends. A tutor should not treat every successful AI interaction as a direct measure of unsupported capability.
Use the traces to generate hypotheses, then confirm consequential conclusions with direct learner performance.
26. AI Engagement and Three-Learner Tuition
In a three-learner group, learner-facing AI can either fragment the room into three isolated screens or create useful differentiated practice.
The tutor should decide which parts of the lesson benefit from shared human discussion and which benefit from individual AI-supported practice. A common explanation can establish the target. Learners can then use different practice levels. The tutor monitors whether the AI support is doing comparable work across learners. A later shared discussion can compare errors or methods without revealing answers prematurely.
Technology should serve the group architecture rather than silently replace it.
27. AI Can Amplify Inequality in Help-Seeking Skill
Fluent learners may ask precise questions and receive better responses. Less confident learners may type vague requests such as “I don’t understand” and receive long explanations they cannot use.
Human support can teach a practical help-seeking structure: state the task, show the attempt, identify the uncertain step, ask for the smallest useful help, then try again.
This is not prompt engineering as an end in itself. The purpose is to help the learner retain ownership of the problem while making support more discriminating.
28. AI Can Make Incorrect Confidence Look Efficient
A learner may receive quick confirmation from the system and move on. If the system is wrong, ambiguous or insufficiently aligned with the intended curriculum, speed becomes dangerous. If the system is correct but the learner did not understand why, confidence can still exceed capability.
The tutor can protect against this with occasional verification: ask the learner to justify a key step, compare with an authoritative source, or solve a fresh item without the answer visible.
Verification frequency should reflect risk. Not every low-stakes practice interaction needs a human audit. High-consequence explanations and unfamiliar content deserve more scrutiny.
29. AI Engagement and Parent Expectations
Parents may understandably ask whether an AI subscription means a learner can practise independently whenever needed. The tutor should explain the implementation question before promising autonomy.
The tool can provide useful practice, but access alone does not tell us whether she will use it well. I am first checking whether she can start the right task, attempt before asking for help and use feedback to correct herself. I will reduce my involvement as those decisions become stable.
Or:
His usage has increased, which is encouraging, but I am not treating minutes as learning. I am checking whether the same method appears on fresh work without the AI open.
This makes the learning criterion visible.
30. AI Engagement and Tutor Workload
Human-supported AI can save tutor time or consume more of it.
If the tutor must continuously interpret the AI, correct it, prompt the learner, choose every task and verify every answer, the tool may not be reducing instructional load. It may simply shift the tutor’s work into supervision.
That can still be worthwhile if the tool creates better practice or differentiation. But the cost should be visible. The tutor should ask which human interventions remain necessary after the learner becomes familiar with the system and whether those interventions are the best use of tutor attention.
31. An Exit Rule Prevents Technology Inertia
Once a family has bought a platform or a programme has adopted it, adults can become reluctant to stop. The tool continues because it exists.
Set a review rule before that inertia develops. If meaningful engagement remains too low despite reasonable support, change the implementation or tool. If engagement is high but later learning receipts remain weak, inspect whether the AI is performing too much of the target operation. If the learner has become independent and the tool remains useful, reduce human supervision.
The success criterion is not loyalty to the technology. It is a better learning route.
32. The Human-Supported AI Engagement Card
- Model: What kind of AI tutoring system is this—learner-facing tutor, adaptive practice, general chatbot, or human-tutor copilot?
- Target: What capability should the learner practise or build?
- Access bottleneck: Is the learner failing to enter or navigate the tool?
- Engagement bottleneck: Is the learner present but passive, answer-seeking or rapidly hint-dependent?
- Human role: Logistical, motivational, interpretive, strategic, instructional or verification?
- Cognitive ownership: Who performs the target operation—the learner, AI or human?
- Meaningful-use receipt: What observable action shows genuine task engagement?
- Learning receipt: What fresh or delayed performance will check whether the capability survives?
- Fade: Which human decisions should the learner eventually take over?
- Exit rule: What evidence would justify changing or stopping the implementation?
33. Do Not Use One AI Study to Justify Another AI Model
Evidence from one AI tutoring system does not automatically transfer to another. Systems differ in content, dialogue design, feedback, guardrails, integration with curriculum, adaptive logic and how much ordinary practice surrounds the AI component.
The NSSA synthesis is useful precisely because it warns against collapsing different models into one category. A tutor should ask whether the research system resembles the system in front of the learner in the aspects that matter to the proposed mechanism.
“AI tutoring research says…” is usually too broad a sentence.
34. Tutor-Facing AI Is a Different Evidence Question
Some systems provide AI support to the human tutor rather than to the learner. For example, the Tutor CoPilot research programme has studied real-time AI suggestions intended to improve human tutoring moves. That model changes tutor decision support, not learner engagement with an AI tutor.
Source: Tutor CoPilot: A Human-AI Approach for Scaling Real-Time Expertise.
The evidence and risks therefore differ. Do not use a tutor-copilot result to claim that learner-facing chatbot use will improve learning, and do not use learner-facing engagement studies to infer effects of a private tutor support tool.
35. The Evidence Base Is Moving Quickly
In June 2026, the Education Endowment Foundation announced a new research programme examining how generative AI affects cognition and learning. The existence of that programme is itself a useful reminder that important educational questions remain unresolved.
Source: Education Endowment Foundation — New research on generative AI and cognition, 8 June 2026.
Tutors should therefore avoid frozen rules such as “AI always helps” or “AI makes learners dependent”. The educational effect depends on the target, design, learner, support conditions and what is measured afterwards.
36. Research Boundary
The 2026 studies cited in this article include working papers. They are useful current evidence, but they have not necessarily completed peer review and should be interpreted with that status visible. Their samples, systems and implementation settings are not interchangeable with private tutoring in Singapore.
Moreover, technology changes quickly. A named AI product can alter its model, interface, feedback design or guardrails after a study is completed. Durable conclusions should therefore focus on mechanisms that can be rechecked: access, actual use, type of engagement, cognitive ownership, feedback action and later performance.
This Handbook volume does not claim that human-supported AI tutoring has one stable average effect. It offers a decision structure for determining what human support is doing in the local learning route and what evidence would justify continuing it.
37. Common Failure Modes
- Access equals intervention: assuming an account or subscription creates learning opportunity automatically.
- Minutes equal learning: treating increased screen time as evidence of capability growth.
- Message-count engagement: rewarding interaction volume without inspecting the learning job.
- Answer-seeking disguised as tutoring: allowing the AI to perform the target operation repeatedly.
- Human babysitting: keeping a tutor beside the learner without a named implementation bottleneck.
- No fade path: increasing AI use only when a human prompts every decision.
- Platform-effect attribution: crediting the AI layer for gains produced by ordinary practice or other programme components.
- Tool loyalty: continuing an ineffective implementation because the subscription already exists.
- Model collapse: treating learner-facing AI, adaptive practice and tutor-facing copilot systems as the same intervention.
- No independent return: ending evaluation with successful AI-supported work rather than checking what survives outside the interaction.
38. The Thirty-Second Human-Supported AI Engagement Gate
What learning operation should the learner perform, what exact bottleneck is the human solving, who is doing the target cognitive work inside the AI interaction, and what fresh performance will show whether increased use became capability rather than only more activity?
If the tutor can answer those questions, AI engagement becomes an interpretable part of the learning route.
39. The Independence Direction
The mature learner should not need a human to tell them every time how to use an AI tutor.
They learn to decide when the tool is appropriate, attempt before outsourcing, ask for the smallest useful help, check whether an answer is plausible, recognise when an explanation has become too long or irrelevant, and close the tool when direct practice would be better.
Most importantly, they learn to distinguish “the AI helped me produce this” from “I can now do this”. That distinction is not anti-technology. It is what makes technology compatible with genuine independence.
Evidence and Connected Reading
- National Student Support Accelerator — AI Tutoring Is Not a Monolith: What We Actually Know (2026)
- EdWorkingPapers — Access Is Not Enough: Human Support Improves Engagement with AI Tutoring (2026 working paper)
- National Student Support Accelerator — Access Is Not Enough study summary
- EdWorkingPapers — One Click Away: AI Tutoring with Khanmigo in a Two-Year School Experiment (2026 working paper)
- EdWorkingPapers — Tutor CoPilot: A Human-AI Approach for Scaling Real-Time Expertise
- Education Endowment Foundation — New research on generative AI and cognition (8 June 2026)
- The Tutor Handbook — Complete Series Index
Final Compression
AI access is not AI use. AI use is not meaningful engagement. Meaningful engagement is not automatically learning.
Human support can repair the pathway between those stages. It can help the learner start, persist, interpret, choose help wisely and use feedback. But the tutor should know which bottleneck the human is solving and should keep the target cognitive work with the learner whenever that is the goal.
Then close the loop with a fresh or delayed learner performance.
The strongest reason to support a learner’s AI use is not that the learner spends more time with AI. It is that the support creates a better opportunity to learn something the learner can later do without the system doing it for them.
That is the Human-Supported AI Engagement Gate.
That is Tutor Handbook Volume 0166.