The Tutor Handbook · Volume 0069 · Series ID THB-0069
The Tutor Handbook: Complete series index.
Two competent tutors look at the same page.
One says, “The learner understands the concept. The problem is careless execution.”
The other says, “No. The execution errors are symptoms. The learner is still selecting the wrong representation.”
Both can point to evidence.
The parent hears two professional voices and wants a simple answer: who is right?
The learner has a more urgent question: what am I supposed to do next?
Professional disagreement is not a defect. It becomes a defect when the adults have no method for turning disagreement into a better evidence question.
The Evidence Conference is a structured conversation in which two or more adults examine the same learner evidence, separate what was actually observed from what each person inferred, align the criterion they are using, locate the exact source of disagreement, and design the smallest new observation that can reduce the uncertainty.
It is not a meeting in which the most senior tutor wins. It is not a vote. It is not a compromise in which one tutor says “weak”, another says “strong”, and everybody settles on “average”. It is not a parent-management exercise designed to make disagreement disappear.
The purpose is narrower and more useful: make the learner model more accurate before disagreement becomes competing interventions.
Quick answer
When tutors disagree, preserve the learner’s work exactly as it was produced. Ask each person to describe the evidence before explaining it. State the target capability and the criterion being applied. Identify whether the disagreement is about the observation, the standard, the cause, the importance of the error or the next action. Then choose one discriminating task, changed condition or additional sample that would make the competing interpretations produce different predictions.
If the new evidence resolves the issue, update the learner route. If it does not, record the uncertainty honestly and choose a low-cost reversible next step. The aim is not perfect agreement. The aim is that adult disagreement never forces the learner to run two contradictory systems while the adults argue about who has better judgement.
1. What this volume owns
This volume owns disagreement about the interpretation of the same learner evidence. That is different from general governance. The wider question of who has authority to decide belongs to the existing eduKateSG owner How Governance Works | Who Gets to Decide?. This volume begins one step earlier: before authority is exercised, can the professionals understand why they are reading the evidence differently?
It also differs from The Learning Claim, which calibrates how far one tutor’s conclusion may travel. The Evidence Conference is what happens when two reasonable claims compete.
And it differs from a formal marking moderation system. Tutors are not awarding a national qualification. The lesson from moderation is methodological: shared criteria, independent evidence reading and explicit comparison can make professional judgement more consistent and defensible.
2. Disagreement is information
If two tutors disagree about a learner, the instinct may be to remove the disagreement quickly. That can waste useful information.
Disagreement often signals one of five things. The tutors may be looking at different evidence. They may be applying different standards. They may agree on the performance but disagree about its cause. They may value different outcomes: one protects accuracy while the other prioritises independence or speed. Or they may be working on different time horizons: one is solving next week’s school problem while the other is protecting a six-month learning route.
Those are not equivalent conflicts. A good conference names which one is present. If the issue is different evidence, collect a shared sample. If the issue is a different criterion, define the criterion. If the issue is cause, run a discriminating task. If the issue is priority, use the Learning Budget and school timing. If the issue is horizon, specify what decision each horizon is trying to protect.
The disagreement becomes useful once it can be located.
3. Freeze the artefact before the story changes
Start with the actual learner evidence.
Keep the original written answer, working, recording, plan, marked paper or task trace. Note the support conditions that mattered. Do not begin with one tutor’s summary such as “she panicked” or “he guessed”. Those are interpretations.
Suppose a Mathematics learner writes the correct equation, makes an early sign error, continues coherently from the wrong expression, then notices the final answer is implausible but does not recover. The artefact contains several observable facts. One tutor may interpret the pattern as execution fragility. Another may call it checking failure. A third may see time pressure as the dominant cause.
All three should begin from the same preserved sequence. Once memory replaces the artefact, hindsight starts editing the learner’s performance. A vivid mistake becomes larger than it was. A late recovery attempt is forgotten. A support prompt that preceded success disappears from the story.
This is why The Decision Record preserves evidence around route changes. The Evidence Conference applies the same discipline before a shared judgement is formed.
4. Observation first, interpretation second
The first round of the conference should sound almost boring.
“The learner selected method B.”
“The first two lines were correct.”
“The sign changed incorrectly in line three.”
“The learner circled the final answer and wrote ‘too large?’ beside it.”
Only after the descriptive layer is stable should people say, “I think this indicates weak sign control,” or “I think the important evidence is that checking detected the implausibility but recovery failed.”
This distinction reduces a common problem: people argue over interpretations while believing they disagree about facts. Once the evidence is described precisely, the remaining disagreement often becomes smaller and more testable.
It also protects the learner from evaluative labels. “Lazy”, “careless”, “bright”, “weak” and “not confident” compress complex observations into identities. A conference should use those words rarely, if at all. Describe the performance first.
5. Agree what capability is being judged
Two tutors can mark the same answer differently because they are answering different questions.
One asks, “Did the learner obtain a correct answer?”
Another asks, “Did the learner select the method independently?”
A third asks, “Did the learner preserve the method under realistic time pressure?”
These are different constructs. The same performance can look strong against one and weak against another.
Before resolving disagreement, state the target in one sentence. For example: “We are judging whether Beatrice can select evidence that directly supports an inference, not whether her final sentence is stylistically polished.” Or: “We are judging whether Denise can maintain accurate algebraic execution across a representative timed section, not whether she can finish a full paper.”
Once the target is explicit, irrelevant evidence can be moved aside. That does not make it unimportant. It simply stops one conversation from carrying five hidden criteria.
6. Agree the reference standard
Professional judgement needs a reference point.
For school assessment, that may be a published marking guide, syllabus standard or teacher criterion. For a tutoring micro-skill, the standard may be an explicitly defined success condition: select the correct method on four of five fresh mixed items without a method cue; explain the causal link in a changed Science context; produce a paragraph in which evidence directly supports the claim.
Without a shared reference, two tutors may use personal prototypes. Tutor A has taught many top-performing learners and calls the response weak. Tutor B compares it with the learner’s previous baseline and calls it strong. Both descriptions can be defensible, but they answer different questions.
The conference should say which reference matters for the decision now: current school participation, the learner’s prior baseline, an independent mastery criterion, examination readiness, or another declared standard. If two standards matter, keep them separate rather than averaging them into an ambiguous label.
7. Independent reading before discussion
When possible, ask each tutor to inspect the evidence independently before hearing the other’s explanation.
This simple step reduces anchoring. The first confident interpretation otherwise becomes the frame through which the second person sees the work. Formal assessment systems use versions of independent marking and standardisation for this reason: the purpose is not to eliminate judgement but to make consistency visible and disagreement inspectable.
Ofqual’s 2025 national assessments regulation report describes marker training and quality assurance in which pre-marked responses are inserted into live marking so consistency against an agreed standard can be monitored. That high-stakes system is far more formal than tuition, and its agreement statistics should not be imported into tutoring. The useful principle is narrower: professional judgement improves when people share a standard and can compare their decisions against common evidence.
A short tutor conference can borrow that discipline without pretending to be an examination board.
8. Name the disagreement type
Once observations and standards are visible, classify the remaining disagreement.
- Observation disagreement: the adults disagree about what occurred.
- Criterion disagreement: they use different definitions of success.
- Causal disagreement: they agree on the performance but explain it differently.
- Priority disagreement: they agree on the weakness but disagree about whether it deserves action now.
- Support disagreement: they disagree about how much help is legitimate or productive.
- Time-horizon disagreement: one optimises immediate school performance, another longer-term independence.
This classification is not a validated instrument. It is a meeting aid. The value is operational: each disagreement type needs a different resolution method.
Do not ask a causal test to solve a priority dispute. Do not ask a parent vote to solve a criterion dispute. Do not use seniority to settle an observation that can be checked by reopening the learner’s work.
9. Causal disagreement needs different predictions
Suppose one tutor thinks Alicia’s Mathematics errors come from method selection. Another thinks the real problem is execution under time pressure.
Arguing harder will not help. Ask what each explanation predicts.
If method selection is the dominant problem, performance should improve sharply when the method is named while time remains the same. If execution under pressure is dominant, naming the method may not remove the sign and transcription errors; relaxing the clock should help more.
Now design two small tasks that alter one relevant condition at a time. The goal is not laboratory certainty. It is to make the explanations answer to evidence.
This is the same spirit as The Differential: keep multiple causes alive until a changed condition separates them. The Evidence Conference adds a human layer—sometimes different adults are carrying the competing hypotheses.
10. Criterion disagreement needs exemplars, not persuasion
Two tutors agree on every feature of Beatrice’s English answer but disagree on whether it is “sufficient”.
That is a criterion problem.
Use an agreed rubric, strong and borderline examples, current school feedback or a set of de-identified comparison responses where appropriate. Ask what feature moves an answer across the boundary. Is it direct textual evidence? Explanation of the link? Scope? Precision? Language?
Current NSW guidance on consistent teacher judgement, updated in 2026, describes moderation as comparing student work with established standards so educators can build a shared understanding of quality. QCAA’s 2026 quality assurance and moderation trials similarly use calibration and consensus conversations around tasks, marking guides and de-identified student work.
Tutoring does not need formal moderation bureaucracy. It can use the same core discipline: disagreement about quality should return to evidence and criteria, not personality.
11. Priority disagreement needs the Learning Budget
Sometimes both tutors are right about the weakness.
One wants to repair it now. The other wants to wait.
That is not an evidence dispute about whether the weakness exists. It is an allocation dispute.
Return to The Learning Budget and The Change Queue. What is upstream? What is urgent? What is already stable enough for maintenance? What school deadline changes opportunity cost? Which intervention has the highest leverage? What will be displaced if this weakness moves to the front?
Do not continue debating the existence of a problem after the real dispute has become whether the problem deserves the next unit of attention.
12. Support disagreement needs a declared target
One tutor wants to remove a prompt. Another wants to keep it.
The correct answer depends on what the learner is currently supposed to own.
If the target is conceptual understanding and the prompt merely clarifies task wording, keeping it may be reasonable. If the target is independent task interpretation, the same prompt may replace the operation being trained.
Ask whether the support changes access or changes the answer-producing process. Ask what evidence exists under less support. Ask whether the support is being installed, held, faded or retired. The current Configuration Snapshot should make that direction visible.
Support disagreement becomes less personal when the learner’s target and current evidence are explicit.
13. Time-horizon disagreement can produce two correct answers
A learner has a major assessment in ten days. One tutor wants targeted rehearsal of the current answer format. Another wants to rebuild the underlying concept more deeply.
Both recommendations may be educationally sensible. They serve different clocks.
The conference should name the clocks. What protects the next school obligation? What protects durable capability? Can the immediate route remain bounded so it does not become permanent? Can deeper repair resume after the assessment? Does the short-term intervention create a dependency or misconception that will cost more later?
Repair, Alignment and Frontier modes help here. Alignment can legitimately dominate temporarily when a real external demand is near. Repair should return when a foundational gap threatens later learning. Frontier should not consume capacity merely because advanced material is interesting.
Disagreement about time horizon is often resolved by sequencing rather than choosing one philosophy forever.
14. Do not vote
Three adults think the learner is ready. Two do not.
The learner does not become ready by majority.
Voting can be appropriate for governance choices among legitimate preferences. It is weak for empirical questions when additional evidence can be collected. If the disagreement is “does this learner select methods independently?”, design an opportunity that requires independent selection. If the disagreement is “has the gain held after delay?”, wait and retest. If the disagreement is “does the prompt replace planning?”, compare performance with and without it.
Where evidence will remain incomplete, a group may still need a decision. In that case, decide using the declared authority structure and record the uncertainty. But do not turn headcount into proof.
15. Do not average incompatible judgements
One tutor rates a response “8 out of 10”. Another rates it “4 out of 10”. The tempting compromise is 6.
That number may have no meaning.
Perhaps Tutor A rewarded mathematical reasoning while Tutor B penalised notation heavily. Perhaps one judged an untimed learning task and the other imagined examination conditions. Perhaps the scale itself is undefined.
Resolve the criterion first. If two legitimate dimensions matter, report both: reasoning strong, notation unstable. A multidimensional description is often more useful than a synthetic midpoint.
Numbers feel objective because they are precise. Precision without shared meaning is decoration.
16. Do not pull rank unless the issue is genuinely authority
The senior tutor may be more experienced. That experience should improve the explanation of evidence, not replace evidence.
“I have taught for twenty years” can justify why a hypothesis deserves serious consideration. It cannot turn an unsupported causal story into fact. The junior tutor may have the fresher observation. The subject specialist may understand one mechanism better. The regular tutor may know the learner’s history. The parent may know what support occurs at home. The learner may know that a tool was used between lessons.
Authority still matters when someone must decide the route. But an Evidence Conference should first make the basis of disagreement visible. Seniority is most useful when it teaches others how to see the work more accurately, not when it closes the conversation before the work is examined.
17. The learner can be part of the evidence conference
Not every conference needs the learner present. But the learner’s account can resolve hidden conditions.
“I knew the method but I couldn’t remember what ‘depreciates’ meant.”
“I copied that sentence from the model answer before I understood it.”
“I used an AI tool to check the paragraph at home.”
“I stopped checking because I had three minutes left.”
These statements are evidence about conditions, not unquestionable explanations. Learners can misremember or misinterpret too. Still, excluding the learner can leave adults debating a process the learner directly experienced.
The long-term goal is greater learner agency. A mature student can begin to say, “I think the problem is selection, not execution; when the method is named I am accurate.” That is a powerful form of self-regulation because the learner participates in the hypothesis rather than merely receiving adult labels.
18. Parents need one operational route after the conference
Parents should not be asked to choose between rival experts without evidence.
A useful post-conference update sounds like:
We agree that Alicia’s current work contains both a method-selection issue and a timing-related execution issue. We disagree about which is primary. For the next two sessions we are keeping the method-selection condition stable and changing only the time condition. If accuracy improves substantially when the clock is relaxed, timing moves forward as the main repair. If the same wrong methods appear untimed, selection remains primary.
This is better than telling the parent that “the tutors have different styles”. It converts disagreement into a bounded evidence plan.
The parent now knows what not to change at home. That matters. If the family simultaneously adds new worksheets, a new timing routine and extra prompting, the conference’s discriminating test disappears inside the Concurrency Problem.
19. A practical Evidence Conference protocol
- Freeze the sample. Preserve the original work and support conditions.
- Describe independently. Each adult records observable features before discussion.
- Name the target. What capability is being judged?
- Name the reference. Which standard, baseline or performance condition matters?
- State each claim. What does each person think the evidence supports?
- Locate the disagreement. Observation, criterion, cause, priority, support or time horizon?
- Generate predictions. If each explanation is true, what should happen under a changed condition?
- Choose one discriminating check. Prefer the smallest task that can reduce uncertainty.
- Decide provisionally. Choose a reversible next route when uncertainty remains.
- Record the unresolved part. Do not rewrite disagreement into fake consensus.
- Return to evidence. Review after the promised sample arrives.
The protocol is an eduKate tutoring procedure, not a validated assessment moderation instrument. Its purpose is disciplined reasoning, not bureaucratic ceremony.
20. Mathematics case: selection or execution?
Fictional case: Alicia loses marks on mixed algebra questions. Tutor One says she still confuses method families. Tutor Two says she selects correctly but becomes inaccurate once the clock is visible.
They review two papers. Both notice that wrong-method attempts cluster late in timed sections. Tutor One sees this as evidence of fragile discrimination. Tutor Two sees it as a consequence of rushing.
The conference produces two predictions. If discrimination is weak, an untimed mixed set should still contain method-selection errors. If timing is primary, untimed selection should remain strong and errors should rise mainly when time pressure returns.
The next session uses a fresh untimed mixed set with no topic labels and no method cues. Alicia selects correctly on nine of ten items and executes accurately. A later moderate-timing set preserves selection but introduces two sign errors.
The evidence does not prove timing is the only cause. It is now strong enough to change priority: selection moves to maintenance; execution under pressure becomes the active job. The conference succeeded because it produced a better route, not because one tutor was declared the winner.
21. English case: weak inference or weak evidence standard?
Fictional case: Beatrice writes, “The character is nervous because she keeps looking at the door.” Tutor One says the inference is sound but under-explained. Tutor Two says the evidence is too general and the inference itself is weak.
They first define the target: select textual evidence that directly supports a plausible inference and explain the link. They compare the exact passage. A second line—where the character folds an invitation repeatedly and asks whether anyone else has arrived—provides more direct support.
The disagreement is partly criterion-based. Tutor One accepted relevant evidence; Tutor Two required the most discriminating evidence. They agree that the learner’s next job is not generic “better comprehension” but selecting between merely related and directly diagnostic evidence.
The next task presents two plausible evidence choices and asks Beatrice to justify which one carries more inferential weight. Now the conference has converted vague disagreement into a teachable distinction.
22. Science case: concept or language?
Fictional case: Ciara gives scientifically incomplete written explanations. One tutor thinks the mechanism is not understood. Another thinks language production is obscuring a sound causal model.
The tutors independently review a fresh oral explanation, a diagram and a written response. Ciara’s diagram and oral chain are coherent; the written version loses one causal connection when she tries to use formal vocabulary.
The conference therefore separates conceptual and expression criteria. The current learner model becomes: causal understanding is usable under diagram-plus-oral conditions; translation into concise written scientific language remains unstable.
This matters because the interventions differ. More concept explanation would be redundant. The route instead practises reconstructing the mechanism first, then compressing it into the written answer form. Later, the diagram scaffold is faded.
Professional agreement emerged not from persuasion but from collecting evidence in more than one mode.
23. Research foundation: moderation and consistent judgement
Several current public assessment systems make the general value of shared criteria and moderation explicit. NSW Department of Education’s 2026 guidance on consistent teacher judgement states that moderation involves comparing student work with established standards to support consistent judgements. Its related on-balance judgement guidance emphasises using a range of evidence and revising professional judgements as more information becomes available.
QCAA’s 2026 moderation trials use quality assurance, calibration and consensus conversations around assessment tasks, marking guides and de-identified student work. Ofqual’s 2025 report describes marker training and ongoing consistency checks in national assessments. These are formal systems with stakes and procedures unlike private tuition.
The Evidence Conference does not claim that their reliability results transfer to tutoring. It borrows a disciplined idea: when professional judgement matters, make the standard visible, make the evidence shareable, and make disagreement inspectable rather than personal.
24. Research foundation: evidence should remain revisable
AERO’s Monitor Progress guidance treats student responses as information teachers use to check learning and adjust teaching, guidance or feedback. The principle is important for disagreement: a judgement should be strong enough to guide action and provisional enough to change when new evidence arrives.
That is also why one conference should not try to solve the learner permanently. It should resolve the current uncertainty enough to improve the next move. A changed-condition task may show that Tutor One’s explanation is stronger today. Two months later the bottleneck may move.
Professional consistency does not mean permanent sameness of judgement. It means adults use evidence and criteria coherently enough that change in judgement can be explained by change in evidence rather than by who happened to look at the learner.
25. Common failure: discuss the learner without the work
“She is usually careless.”
“He is actually very capable.”
“She panics in tests.”
These may contain useful historical information. None should substitute for the artefact currently under dispute.
Start from this answer, this task, this support condition, this time window. Then bring history in as context and competing hypotheses. Otherwise reputations enter the room before evidence does.
26. Common failure: make consensus the goal
Sometimes the correct conference outcome is:
We agree on the observation and the target. We still disagree about whether timing or selection is the primary cause. We are running one untimed mixed check and one moderately timed parallel check before changing the route.
That is a successful conference.
Fake consensus is dangerous because it erases uncertainty without resolving it. The next tutor then inherits a confident learner model that nobody truly believed.
Preserve unresolved disagreement when it matters. Make the next evidence event responsible for reducing it.
27. Common failure: conference every tiny disagreement
Professional collaboration has a cost. Do not convene a full Evidence Conference because two tutors prefer different examples or because one would phrase feedback differently.
Use the process when the disagreement could materially change diagnosis, support, route, workload, readiness, release, re-entry or a consequential claim about the learner.
Minor differences can remain local. The system needs enough consistency to protect the learner, not uniformity of every teaching move.
28. The Evidence Conference card
- Artefact: What exact learner work are we discussing?
- Conditions: What support, time and resources were present?
- Independent observations: What did each adult notice before hearing the other?
- Target: What capability are we judging?
- Reference: What standard, baseline or school requirement matters?
- Claim A: What does the first interpretation say?
- Claim B: What does the second interpretation say?
- Disagreement type: Observation, criterion, cause, priority, support or horizon?
- Prediction: What should happen if each claim is right?
- Discriminating check: What small new task can separate them?
- Provisional route: What do we do while uncertainty remains?
- Review trigger: What evidence brings us back together?
29. The ethical standard
Adult disagreement should not become learner instability.
One tutor should not tell the learner to slow down while another tells the learner speed is the priority. One adult should not restore a scaffold while another is trying to fade it. One person should not describe a weakness as repaired while another quietly runs full repair every week.
Disagreement is allowed. Competing live systems are not a harmless consequence.
When adults disagree about learner evidence, the learner should receive one coherent provisional route and a clear promise about what new evidence will make the adults reconsider it. Professional judgement earns trust by remaining answerable to the work, not by becoming louder.
Evidence and connected reading
- The Tutor Handbook Vol No.0068 | The Opportunity Check
- The Tutor Handbook Vol No.0049 | The Learning Claim
- The Tutor Handbook Vol No.0036 | The Differential
- The Tutor Handbook Vol No.0064 | The Change Queue
- eduKateSG | How Governance Works
- NSW Department of Education | Consistent Teacher Judgement
- NSW Department of Education | Forming an On-Balance Judgement
- Queensland Curriculum and Assessment Authority | Quality Assurance and Moderation Trials
- Ofqual | National Assessments Regulation Annual Report 2025
- Australian Education Research Organisation | Monitor Progress
Final compression
Keep the learner’s work.
Describe it before explaining it.
Name the target.
Name the standard.
Locate the disagreement.
Make the competing explanations predict different things.
Collect the smallest new evidence that can separate them.
Do not vote. Do not average incompatible claims. Do not pull rank when the evidence can still answer.
The Evidence Conference turns professional disagreement from a contest of confidence into a design problem: what do we need to observe next so the learner receives a better route?
That is the Evidence Conference.
That is Tutor Handbook Volume 0069.