The Tutor Handbook · Volume 0111 · Series ID THB-0111
Two months into tuition, a parent asks the question every serious tutor should expect: “How much has my child improved?”
The tutor has plenty of evidence from now. The learner starts work more independently. Several error patterns have reduced. Recent school results are stronger. The tutor can point to cleaner explanations and better checking. There is one problem: nobody collected a clean starting measure before teaching began.
The temptation is to rebuild the past. The tutor remembers that the learner was “very weak” at the start. The parent remembers repeated homework struggles. There is an old school paper with a low mark, but it covered different content under different conditions. The first tuition lesson contains work samples, yet the tutor explained several items before the student attempted them. It feels obvious that progress happened. It may well have happened. But the exact size of that progress is not recoverable merely because the story is plausible.
The direct answer is: do not invent a baseline after the fact. Separate what can be supported from what cannot. Use surviving early evidence descriptively, state its limitations, establish a clean current baseline for the capability that matters now, and monitor forward from there. A tutor can honestly say, “Current evidence is stronger than the early work we still have, but we did not collect a comparable baseline before teaching, so I cannot quantify the exact amount of improvement attributable to tuition.” That sentence is not weakness. It is measurement discipline.
A baseline is a starting condition, not a memory
In ordinary tutoring language, “baseline” often means “what the child was like at the start”. In stronger evidence language, a baseline is more specific. It is an observed starting level on a defined target, collected under conditions that allow later results to be interpreted against it. The target might be reading fluency, accurate algebraic manipulation, evidence selection in comprehension, retrieval of key vocabulary, planning a response or another capability. The baseline should match the job we later want to discuss.
A memory such as “she used to need a lot of help” can still be useful history. It is not automatically a clean measure. An old school mark can still be relevant. It may not be comparable to a new tuition task. A first-lesson page can be valuable. If the tutor coached half the answers, it cannot be treated as unaided pre-intervention performance. The Missing Baseline begins by protecting these distinctions.
Why this problem is common in real tutoring
Tutors often begin teaching because a learner needs help now, not because a research protocol is waiting. A child arrives with homework due tomorrow. A school test is close. A parent has brought a marked paper and wants the first weak link repaired. The tutor explains, teaches and stabilises. Only later does somebody ask for a precise account of change.
This is understandable. It also means the starting evidence may already be contaminated by intervention. Once the tutor has taught the concept, supplied the method, changed the learner’s strategy or repaired a prerequisite, a true pre-teaching baseline for that exact condition no longer exists. The tutor can create a current baseline. They cannot travel backwards and observe the untouched starting state.
The correct response is not guilt and not statistical theatre. It is to narrow the claim.
Three questions that must be kept separate
Parents and tutors often compress three questions into one sentence: “Did tuition work?” The evidence may support these questions differently.
- What can the learner do now? This can often be answered well with current evidence.
- How different is that from the learner’s starting point? This requires comparable earlier evidence and may be harder.
- How much of the change was caused by tuition? This is a causal claim and is harder still because school teaching, home study, maturation, practice, examinations and other changes may have contributed.
A missing baseline mainly damages the second question, but it also makes causal storytelling easier to overstate. The tutor can still teach effectively without solving causal inference. The educational obligation is to say only what the evidence supports.
Do not turn an old mark into a fake baseline
Suppose the learner scored 48 on a school paper before tuition and 72 on a different school paper eight weeks later. The arithmetic difference is 24 marks. It is tempting to report “a 24-mark improvement”. As a description of two scores, that is true. As a measure of capability change, it may be much less secure.
The papers may differ in topic coverage, difficulty, marking, time pressure and opportunity to learn. The learner may have encountered the second content more recently. The first paper may have been unusually poor for reasons unrelated to stable knowledge. The second may contain familiar question types. School instruction has continued during the same period. The tutor should not erase those differences because the numbers are convenient.
A better report is: “The school score rose from 48 to 72 across two different papers. That is encouraging external evidence, but because the papers are not a controlled matched measure, I am not treating the 24-point difference as a precise estimate of tuition effect.” The parent still receives useful information. The tutor avoids pretending that unlike assessments are interchangeable.
Do not turn recollection into a number
Human memory is good at stories and poor at recreating exact measurement conditions. After a learner improves, the starting point can look worse in retrospect. After a learner struggles, early strengths can be forgotten. Tutors also remember dramatic errors more easily than ordinary successful attempts. Parents remember emotionally expensive evenings. Learners remember embarrassment or relief. These recollections matter as experience, but they should not be converted into pseudo-precision.
Do not write “baseline 30%” because “that feels about right”. Do not assign a 1-to-5 independence score retrospectively unless the scale and observations existed at the time. Do not reconstruct prompt frequency from memory as if every lesson had been coded. If the evidence is qualitative, keep the claim qualitative.
The Evidence Salvage Table
When the clean baseline is missing, early evidence can still be salvaged without pretending it is better than it is. Sort it by what it can legitimately show.
- Direct early work sample: shows what was produced on that task under those recorded conditions.
- Marked school assessment: shows performance on that assessment, subject to its content and conditions.
- Tutor note: shows what the tutor observed or inferred at the time; stronger when observation and interpretation were separated.
- Parent report: provides contextual information about home study, task completion or recurring concerns, but is not a standardised measure.
- Learner report: provides valuable perception of difficulty, confidence and strategy, but should not be treated as objective performance data by itself.
- Memory reconstructed later: may guide questions, but should be labelled retrospective and not used to manufacture numeric precision.
The purpose is not to rank people’s credibility. It is to preserve evidence provenance. A parent’s observation that homework once took three hours may be highly relevant to family burden. It still answers a different question from an independently completed comprehension probe.
Composite case: Alicia and the old Mathematics paper
The following case is fictional and composite. Alicia arrives for tuition after scoring 41 on a Mathematics paper. Her tutor begins repair immediately. Eight weeks later she scores 68 on another school paper. The family is delighted and asks how much the tuition improved her Mathematics.
The tutor first compares the actual papers. The earlier paper included several topics Alicia had not fully studied. The later paper emphasised topics recently taught in school and tuition. The first was completed during a week when Alicia had missed school; the second was not. The later paper also had a different distribution of marks. The two scores remain useful milestones, but they are not a clean before-and-after pair.
The tutor therefore adds current capability evidence. On a fresh mixed set that samples the repaired algebra skills, Alicia selects methods without chapter labels, shows working and checks her solutions. A week later she performs similarly on changed numbers and wording. The tutor can now say: “Alicia’s current algebraic control is materially stronger than what is visible in the early work we have. Her school results also improved. Because we did not collect a comparable independent baseline before teaching, I cannot assign an exact tuition-produced gain.”
The statement is precise where precision is earned and restrained where it is not.
Composite case: Beatrice and the first lesson that became teaching too quickly
Fictional composite case. Beatrice begins tuition with difficulty in Science explanations. During the first lesson, the tutor asks one question, sees a weak answer and immediately teaches an observation-concept-mechanism structure. Beatrice then practises with prompts. Two months later her explanations are far stronger.
The original weak answer is real evidence. It is not a comprehensive baseline for the whole explanatory skill. One item may have been unusually difficult. The tutor did not sample multiple topics, delay or changed contexts before teaching. The first lesson therefore supports a narrow claim: Beatrice initially produced an incomplete explanation on that task before instruction. It does not support “Beatrice began at 20% mastery”.
The tutor now establishes a current baseline across several fresh explanation tasks under recorded support conditions. Future growth can be monitored from this point. The earlier sample stays in the narrative as an anchor, but its evidential role remains bounded.
Composite case: Ciara and the parent’s memory
Fictional composite case. Ciara’s parent remembers that homework used to require constant supervision. After a term of tuition, Ciara starts more tasks alone and asks more specific questions. There is no formal starting count of prompts or time-on-task.
The parent report should not be discarded simply because it is not standardised. It describes a meaningful family outcome. The tutor can report it as such: “Parent and learner reports indicate that task initiation at home has become less dependent on continuous adult prompting.” The tutor should not add, “prompt dependence decreased by 60%” unless somebody actually collected the data needed for that number.
From now on, a simple prospective measure can be introduced if useful: one brief weekly note on whether Ciara began a defined homework task independently, after one reminder or after sustained prompting. The measure starts now. It does not retroactively quantify the past.
Establish a current baseline, not a retroactive one
Once the problem is noticed, the tutor can create a clean current starting point for future monitoring. Define the target narrowly, choose tasks that actually sample it, record relevant conditions and collect enough evidence to avoid treating one unusual attempt as the whole learner.
For some well-defined skills, multiple short probes may be useful. University-based progress-monitoring resources from the IRIS Center, for example, describe establishing a stable baseline using several data points over a short period and often summarising them with a median. That guidance belongs to structured school progress-monitoring contexts; a private tutor should not mechanically transplant its numbers to every subject or capability. The transferable principle is that a baseline should be based on observed performance and should match what will later be monitored.
For complex writing, reasoning or problem solving, a single numeric score may hide more than it reveals. A current baseline may need several dimensions: task interpretation, method selection, execution, explanation quality, checking and support level. The goal is not to create a complicated dashboard. It is to preserve the aspects of performance that the future claim depends on.
The forward line begins today
A missing baseline is frustrating because it limits the story about yesterday. It does not prevent better evidence tomorrow. Mark the date when systematic monitoring begins. From that point, keep the target and conditions sufficiently stable to make comparison useful, while allowing teaching to adapt between checks.
This distinction matters: teaching should not become frozen merely to protect measurement. A tutor exists to help the learner, not to preserve a perfect experiment. Monitor at planned moments, teach responsively between them, and record major route changes that affect interpretation. The tutor can improve evidence discipline without sacrificing instruction.
Baseline does not mean “lowest score”
A baseline should not be chosen because it makes later progress look impressive. If the learner produced 42, 61 and 58 on three reasonably comparable starting probes, selecting 42 as “the baseline” exaggerates the apparent gain. Likewise, selecting the best early performance can understate difficulty. Where several comparable observations exist, use a defensible summary rather than the most dramatic point.
The exact summary depends on the measure. In structured curriculum-based progress monitoring, median baselines are common because they reduce the influence of one unusual point. In ordinary tutoring, the more important discipline is transparency: say what observations existed, how they were summarised and whether the tasks were genuinely comparable enough to justify a numeric trend.
Baseline does not prove causation
Even a good before-and-after baseline does not automatically prove that tuition caused the difference. The learner is also attending school, doing homework, maturing, revising independently, receiving parent support and encountering new curriculum content. Some changes occur because of practice with the measure itself. Some apparent gains reflect easier tasks. Some genuine tuition effects interact with school teaching and cannot be cleanly separated.
The Tutor Handbook’s existing Learning Claim already protects the distinction between progress evidence and causal attribution. The Missing Baseline applies that discipline to the common case where the tutor cannot even make a strong quantitative before-after comparison.
What a tutor can say with confidence
Strong language is not necessarily numerical language. A tutor can make useful statements such as:
- “Current independent work shows stable use of the repaired method across three varied tasks.”
- “The earliest surviving work contains repeated denominator errors; those errors have not appeared in the last four comparable checks.”
- “The learner now begins this task without the prompt that was still required six weeks ago.”
- “School results have improved, but the papers differ, so I am treating them as external confirmation rather than a precise matched measure.”
- “We did not collect a clean starting baseline, so I cannot quantify the exact gain from the first day.”
These sentences give parents something more valuable than an invented percentage: a map of what changed, under what conditions, and how strong the evidence is.
What a tutor should not say
- “Tuition improved her by 35%” when no comparable baseline exists.
- “She was at Level 2 when she started” when the level was assigned retrospectively.
- “The old school mark is her baseline” without checking task comparability.
- “We know the tuition caused the improvement” when several systems changed together.
- “No baseline means we have no evidence at all.”
The last error matters. Missing one ideal measure does not make every other observation worthless. It changes the size of the claim that can be defended.
A baseline can be qualitative when the target is qualitative
Some tutoring targets resist useful compression into one score. Consider the quality of an argumentative paragraph. The tutor may care about claim precision, evidence relevance, explanation of the evidence, counterargument handling and coherence. A single mark can summarise these, but it may not reveal what changed.
A qualitative baseline can therefore preserve a structured sample: one or two pieces of work, the support level, the recurring break and an annotated description of what the learner can currently do. Later samples can be compared against the same dimensions. The important constraint is consistency of interpretation. Qualitative does not mean vague.
Avoid baseline inflation through over-testing
Once tutors discover the value of baselines, they can overcorrect. Every topic begins with a battery of tests. Lessons become data collection. Learners feel that tuition is an endless assessment centre. The tutor has more numbers and less teaching time.
The solution is proportionality. Collect baseline evidence when it changes a decision, supports a meaningful progress claim or protects against misdiagnosis. Use naturally occurring work when it is sufficiently clean. Keep probes short when the target allows it. Do not measure a capability more precisely than the educational decision requires.
The National Student Support Accelerator’s tutoring standards emphasise formative assessment and student progress, but the purpose is to improve tutoring, not to create a miniature research institution around every child. Evidence should serve learning.
The baseline and the three tuition modes
Repair, Alignment and Frontier remain the three tuition modes. The need for a baseline changes slightly across them. In Repair, a baseline can document the weak link and later show whether the repair survives. In Alignment, it can show whether the learner has become more stable relative to current school demands. In Frontier, it can prevent harder work from being mistaken for regression merely because the challenge level increased.
A missing baseline does not require a fourth mode. It requires a measurement boundary. Establish where reliable monitoring begins and avoid rewriting the earlier story with false precision.
Where the Tutor Classification Model helps
The Tutor Classification Model matters only where the tutor’s function changes the evidence need. A Class 3 Diagnostic Tutor benefits from a clean current baseline when deciding whether an intervention should target one weak link or another. A Class 4 Route Designer needs enough baseline information to know whether a route change produced the expected shift. A Class 5 Performance Coach may need a separate baseline under realistic performance conditions. A Class 6 Learning Architect must avoid combining unlike evidence streams into a single story. These are functions, not permanent rankings.
Fresh research: progress monitoring
The Australian Education Research Organisation’s current Monitor progress guide emphasises checking student understanding, identifying gaps and using evidence to adjust teaching and feedback. It is research-informed school guidance rather than a private-tuition baseline protocol, but it supports the principle that progress decisions need observable evidence rather than impression alone.
The IRIS Center at Vanderbilt University provides detailed educational progress-monitoring materials. Its guidance on creating a goal line describes establishing a baseline from several data points collected over a short period and using the median as a stable starting estimate in the curriculum-based measurement context. A related progress-monitoring module similarly begins with current performance before setting goals and collecting repeated measures.
Those procedures should not be copied blindly into every tutorial. They are strongest when the measure is designed for repeated progress monitoring and the target is sufficiently defined. A complex essay, a novel problem-solving task and a one-minute fluency probe do not produce the same kind of data. The general lesson is to establish the starting point with the same kind of evidence that will later be interpreted as progress.
Fresh research: tutoring evidence and implementation
Stanford’s National Student Support Accelerator includes formative assessment and student progress in its Tutoring Quality Standards. It distinguishes research-based, research-informed and emergent standards rather than treating every programme feature as equally established. That evidence taxonomy is useful for this article because baseline discipline should not be used to manufacture certainty that tutoring research itself does not provide.
The Education Endowment Foundation’s Effective tutoring summary reports positive average impacts for tutoring while noting variation and the importance of monitoring and implementation. Its implementation guidance also emphasises understanding what an approach looks like in day-to-day practice. These sources support disciplined monitoring, but they do not justify attributing a particular learner’s entire improvement to tuition from a simple before-after score.
The first clean baseline may arrive after improvement has already begun
This can feel disappointing. If the learner has already improved, a current baseline starts too high to capture the earliest gain. That lost measurement opportunity is real. Do not compensate by lowering the baseline artificially.
Instead, divide the story. The early phase is a bounded historical phase: describe surviving evidence and its limitations. The later phase is a prospective monitoring phase: establish comparable measures and track forward. Over time the second phase becomes increasingly useful, and the need to reconstruct the first phase diminishes.
A handoff should preserve the baseline boundary
If a tutor changes, the new tutor should know where clean monitoring begins. Otherwise the next person may treat a later current baseline as though it were the learner’s original starting point, or treat an old coached worksheet as independent evidence. A one-line note can prevent the error: “Comparable independent monitoring begins from 13 September; earlier evidence is descriptive and comes from mixed conditions.”
This protects continuity without rewriting history. It also helps a parent understand why the centre can be precise about recent progress while more cautious about the first weeks.
The delayed and changed-condition receipt
Once the current baseline exists, future evidence should not remain trapped in near-duplicate tasks. If the capability is meant to transfer, check it after delay and under a meaningfully changed surface. Keep conditions comparable enough that the change is interpretable. Record support level. If the learner succeeds only on immediate near copies, the progress claim should remain modest. If the learner succeeds later, selects the route independently and survives variation, the claim can grow.
The new monitoring system should therefore avoid the opposite problem of hyper-standardisation. Repeated identical items can create practice effects. Good evidence needs enough consistency to compare and enough variation to test whether the learning travels.
A parent-facing script without false precision
A calm explanation can be direct: “I can show you several ways your child’s current work is stronger, and the recent school result is encouraging. I cannot honestly give you an exact percentage improvement from the first lesson because we started teaching before collecting a comparable baseline. From this month, I have a cleaner set of measures, so future progress claims will be tighter.”
That answer respects the parent’s question. It does not hide behind technical language, and it does not manufacture certainty to sound professional. The tutor is saying where the evidence is strong, where it is weak and what has changed in the monitoring process.
A compact Missing Baseline protocol
- Define the exact progress question now being asked.
- Locate all surviving early evidence.
- Label each source by its real conditions and limitations.
- Do not convert memories or unlike scores into a fabricated starting number.
- Separate current capability, change from starting point and causal attribution.
- Create a clean current baseline for the target that matters now.
- Use comparable repeated evidence prospectively.
- Mark major route or condition changes that break comparability.
- Check learning after delay and meaningful variation.
- Report only the size and cause of change that the evidence can actually support.
Failure mode: baseline theatre
Baseline theatre happens when a tutor creates an impressive-looking measurement system after the fact. Old work is rescored using a new rubric that did not exist. Memory is converted into ratings. Different assessments are normalised informally. A chart is drawn. The resulting line looks scientific because it has axes and numbers.
Nothing about a graph repairs weak inputs. A tidy formula is not a validated measure. If the historical evidence was mixed, the graph simply hides the mixture behind precision. Better to show a shorter, honest timeline: “Early evidence mixed and partly supported → clean baseline starts here → repeated comparable checks from here.”
Failure mode: refusing to recognise progress because the baseline is imperfect
Evidence discipline can become too severe. A tutor may say, “We have no baseline, so we cannot say anything about progress.” That ignores useful evidence. A learner who once could not begin any mixed algebra question without a method prompt and now independently selects and completes several changed problems has changed in an educationally meaningful way, even if the exact initial rate was not measured.
The correct language is proportional. Strong qualitative change can be reported qualitatively. Repeated current evidence can establish present capability. School returns can corroborate. Parent and learner reports can describe functional change. What should be withheld is unsupported precision and unsupported causation, not all recognition of improvement.
The final return
A missing baseline is not a reason to invent a past. It is a reason to draw a line in the present.
Before that line, preserve the evidence you really have: old work, school results, parent observations, tutor notes, learner recollections. Keep each in its proper category. Do not force them into one number. After that line, monitor the target deliberately enough that future questions about progress have better answers.
The tutor’s job is not to win the argument that tuition worked. The tutor’s job is to improve learning and describe that improvement as truthfully as the evidence allows. When the starting measure is missing, truthfulness means accepting one uncomfortable sentence: we may know that the learner is stronger now without knowing exactly how many units of change belong between then and now.
That is still useful knowledge. From today onward, it can become better measured knowledge.
Sources and evidence boundary
- Australian Education Research Organisation — Monitor progress: current research-informed guidance on checking student learning and adjusting teaching.
- IRIS Center, Vanderbilt University — Create a Goal Line: structured progress-monitoring guidance describing baseline collection from multiple data points in curriculum-based measurement contexts.
- IRIS Center — Progress Monitoring: guidance on collecting current baseline data before setting goals and monitoring intervention response.
- Stanford National Student Support Accelerator — Tutoring Quality Standards: tutoring-specific evidence framework including formative assessment and student progress.
- Education Endowment Foundation — Effective tutoring: tutoring evidence summary showing positive average effects alongside variation and implementation cautions.
- Education Endowment Foundation — Implementation guidance: research-informed guidance on monitoring how an approach operates in real settings.
The Missing Baseline is an eduKate tutoring reasoning framework, not a validated scale, causal estimator or official school assessment procedure. Some cited progress-monitoring methods were developed for structured educational measures and cannot be transferred automatically to complex tuition outcomes. All worked cases are fictional composites. The article supports bounded educational reporting, not clinical diagnosis or guarantees of tutoring effect.