The Tutor Handbook · Volume 0206 · Series ID THB-0206
The Tutor Handbook: Complete Series Index
“What does a good answer look like?” can be either a learning question or an answer leak
A learner is told to improve a Science explanation.
The tutor says, “Include the cause, the process and the outcome.”
The next answer is much better.
Did the learner learn how to construct a scientific explanation?
Perhaps.
Or perhaps the tutor supplied the hidden structure that the learner was supposed to select independently.
In another lesson, a learner is asked to write a situational response. The tutor hands over a checklist containing every required content point. The learner ticks each item and produces a complete answer.
The checklist is useful.
It may also make the task easier in exactly the place the learner needs to become independent.
Success criteria matter because learners need to understand what they are trying to achieve. Vague goals such as “write better”, “be more accurate” or “show your working” are difficult to act on.
But success criteria can become over-scaffolding.
The Success-Criteria Transparency Gate asks how much of the quality standard should be made visible before, during and after a task so the learner can monitor progress without having the task’s key decisions performed for them.
The gate protects two truths.
Learners should not be expected to hit a target that has never been made intelligible.
Learners should also not be told so much about the target structure that success becomes compliance with the tutor’s template rather than independent performance.
Quick answer
Make the learning goal clear.
Make the quality standard clear enough that the learner can understand what counts as successful performance.
Then decide which parts of the success criteria are legitimate to reveal before the task.
If a criterion describes the enduring quality of the target—accuracy, causal coherence, justification, relevance, complete reasoning, audience fit—it can often be shared.
If a criterion names the exact content, method, evidence or step that the learner is supposed to generate, revealing it may change the task.
Use examples and contrasts to make quality visible without reducing it to a recipe.
Let criteria become more learner-owned over time.
Early in learning, the tutor may provide explicit criteria.
Later, ask the learner to recall, select, explain or generate the criteria.
Before final verification, remove criteria that would supply the target decisions.
After the task, use the criteria again for feedback, self-checking and revision.
The educational objective is not to hide standards.
It is to make standards visible while preserving the intellectual work that the learner must eventually perform alone.
The ownership boundary
This gate sits near several existing owners.
The Task-Purpose Gate asks whether a task is for teaching, practice, diagnosis, progress monitoring or verification before deciding what support is allowed.
How Feedback Works in Teaching owns the broader problem of turning evidence into a better next attempt without replacing learner thinking.
The Criterion Drift Check asks whether the judgement standard has quietly become harsher, softer or otherwise different over time.
The Response-Modality Equivalence Gate asks whether different response routes still measure the same target.
The present article owns what happens before and around performance when the tutor decides how visible the success standard should be.
The criteria are part of the task environment.
They can support self-regulation.
They can also become cues.
What current guidance says
The Australian Education Research Organisation’s Explain learning objectives, published in February 2024, describes learning objectives and success criteria as important parts of explicit instruction and formative assessment. It emphasises clear, measurable objectives, success criteria and connections to prior knowledge.
AERO’s Formative assessment practice guide recommends setting clear and measurable learning objectives, clearly explaining success criteria, and making the purpose and relevance of tasks visible.
AERO’s Explain learning objectives: Primary and secondary describes practice in which learning objectives and success criteria are explained, referred to through the lesson and used to check achievement.
The National Student Support Accelerator’s current Tutoring Session Structure recommends framing sessions with a learning objective and aligning formative assessment to what was modelled and practised.
EEF’s Embedding Formative Assessment identifies clarifying and sharing learning intentions and success criteria as one of the five key formative-assessment strategies in the programme.
EEF’s Feedback, published 24 August 2026, frames useful feedback relative to goals or outcomes and emphasises helping learners understand what success looks like and what next steps move learning forward.
These sources support clarity of learning goals and success. They do not imply that every criterion should be given in full before every task, or that a checklist is always the best way to express quality.
That is the tutoring decision this gate owns.
Separate the learning goal from the task instructions
“Complete questions 1–10” is not a learning goal.
“Use evidence to justify an inference” can be.
“Write three paragraphs” is not automatically a success criterion.
“Organise the explanation so the cause leads logically to the outcome” can be.
Tutoring becomes more coherent when the tutor can state the learning target without naming the activity.
A task is the vehicle.
A criterion describes successful performance on the target.
Confusing these creates shallow transparency.
The learner knows what to do physically but not what they are meant to get better at.
A useful test is to ask:
“If I changed the worksheet, would this learning goal still make sense?”
If yes, the tutor is probably naming a durable learning object rather than an exercise instruction.
Success criteria can describe quality at different resolutions
A criterion can be broad.
“The explanation is scientifically accurate.”
It can be more operational.
“The explanation links the relevant condition to the mechanism and the outcome.”
It can become highly specific.
“State that the air near the cold surface cools, water vapour condenses and liquid droplets form.”
The third version may be a model answer disguised as a criterion.
Resolution should match the task purpose.
During early teaching, high-resolution criteria can help make hidden structure visible.
During supported practice, the tutor may retain the structure but remove exact content.
During verification, the learner may need to generate the relevant criteria or perform without the list.
The criteria should fade as the learner becomes capable of carrying the standard internally.
Composite case: the Science checklist that writes half the answer
This case is fictional and constructed for teaching.
Alicia struggles to answer open-ended Science questions.
Her tutor gives a checklist:
1. State the condition.
2. Name the process.
3. Link the process to the outcome.
4. Use the key vocabulary.
Alicia’s next responses improve.
The tutor could conclude that she now understands how to answer.
Instead, the tutor tests the support.
On a fresh question, Alicia receives only two criteria:
“Accurate mechanism.”
“Complete causal link.”
She hesitates but produces a workable answer.
On another fresh question, the criteria are removed. Alicia identifies the condition and outcome but omits the mechanism.
The original checklist was useful teaching support. It was not yet proof of independent explanatory control.
The tutor now knows which part has become internal and which part still depends on the visible structure.
The checklist has done its job because it can begin to disappear.
Criteria should name quality without supplying the content when possible
Compare two criteria for an English inference question.
“Use one relevant piece of evidence and explain how it supports your inference.”
Versus:
“Mention that the character avoids eye contact and therefore feels guilty.”
The first names a quality structure.
The second supplies the evidence and inference.
Both can be useful in teaching.
They are not equivalent evidence conditions.
Before an independent check, prefer criteria that describe what a strong response must accomplish rather than what the learner must specifically say.
This preserves target generation.
Examples can make quality visible better than abstract criteria alone
Learners may not understand “coherent reasoning” until they see what coherent and incoherent reasoning look like.
Use contrasting examples.
One answer has correct facts but no link.
One has a clear link but irrelevant evidence.
One is complete and precise.
Ask the learner to compare.
“What makes this one stronger?”
“Which criterion does this sentence satisfy?”
“What is missing here?”
Examples can give criteria meaning.
Then remove the example.
The learner should eventually be able to recognise the quality without the model sitting beside the task.
The Erroneous-Example Gate remains relevant when weak or incorrect examples are deliberately used for comparison.
Do not turn criteria into a rubric the learner cannot operate
A detailed rubric can be educationally sophisticated and practically useless.
The learner reads twelve rows, four levels and dense descriptors.
Working memory is now spent on decoding the rubric.
Criteria should be usable during the task.
A small set of high-value dimensions often works better.
For a paragraph:
Claim.
Evidence.
Link.
For a multi-step Mathematics solution:
Correct method.
Valid working.
Logical sequence.
Final answer appropriate to the question.
For a Science explanation:
Relevant condition.
Accurate mechanism.
Causal connection.
The exact dimensions depend on the task.
The principle is economy.
A criterion that cannot guide a learner’s next action is documentation, not necessarily instruction.
Success criteria should not become a substitute for judgement
Checklists create a tempting fiction.
All boxes ticked, therefore success.
But quality is often relational.
A paragraph can contain evidence and still use irrelevant evidence.
A Science explanation can contain a process word and still reverse cause and effect.
A Mathematics solution can show working and still choose a method that does not fit the condition.
A situational response can include every content point and still fail audience or purpose.
The learner needs judgement, not only item completion.
A good criterion invites a quality decision.
“Is the evidence relevant to the claim?”
Not merely:
“Did I include evidence?”
That shift is small and important.
Composite case: the writing checklist that produces formulaic paragraphs
This case is fictional.
Beatrice learns a paragraph structure with a visible checklist.
Topic sentence.
Evidence.
Explanation.
Link.
Her paragraphs become complete.
After several weeks, every paragraph sounds the same. She forces a “link” sentence even when the paragraph is already complete. She chooses evidence because one piece is required rather than because it is the strongest support.
The tutor realises the checklist has moved from scaffold to constraint.
The next stage changes the criteria.
“Make one clear claim.”
“Use the strongest relevant support.”
“Explain the relationship.”
The fixed four-part structure becomes optional.
Beatrice must now decide how to realise the quality.
The criteria become less procedural and more disciplinary.
This is a common progression.
Early criteria can be concrete.
Later criteria should preserve standards while reopening judgement.
Learners should progressively generate the criteria
One of the strongest signs that criteria are becoming internal is that the learner can state them without being shown the list.
Before a task, ask:
“What will make this answer strong?”
“What are you going to check before you finish?”
“What does the marker need to see?”
“What would make this method valid?”
The learner’s answer is evidence.
If they can generate the quality dimensions, self-monitoring becomes possible.
If they cannot, the tutor may need to reteach the standard before asking for independent regulation.
The goal is not to memorise the tutor’s wording.
The goal is to carry a useful model of quality.
Criteria can create cue dependency
A learner becomes accurate only when the checklist is visible.
The checklist may now be performing part of the monitoring function.
That is not a reason to remove it immediately.
It is a reason to fade it deliberately.
Full criteria.
Shortened criteria.
Criteria recalled from memory.
Criteria generated by the learner.
No visible criteria during the attempt.
Criteria returned afterward for self-review.
Fresh task later.
This sequence turns external quality control into internal monitoring.
The Independence Test owns the broader judgement of whether self-regulation has become stable.
Task purpose changes how much transparency is legitimate
During teaching, explicit criteria can reveal hidden structure.
During practice, criteria can support deliberate attention.
During diagnosis, criteria may need to be reduced if the tutor is trying to discover whether the learner can identify the relevant dimensions independently.
During progress monitoring, criteria should match the evidence claim. If the same checklist has always been visible, the tutor should not suddenly interpret success as unsupported performance.
During verification, criteria that supply the target decisions may need to disappear.
There is no contradiction between transparent teaching and independent assessment.
The support condition changes because the decision changes.
The target can contain criteria that should never be hidden
Some standards are not clues.
They are the definition of the task.
An essay is being written for a particular audience.
A Mathematics proof requires logical justification.
A Science investigation requires a fair test.
A situational response has a specified purpose.
A student should know these.
Hiding the task definition to make the work “harder” is not rigorous.
The gate is about criteria that supply the solution route, not criteria that state the legitimate construct.
Ask:
“If I reveal this criterion, am I clarifying what counts as quality—or am I revealing the decision the learner should make?”
That question catches many boundary errors.
Criteria should align with the actual curriculum or assessment target
Tutors sometimes invent criteria that reward their preferred style rather than the real learning goal.
“Always use three examples.”
“Always write five paragraphs.”
“Always draw a bar model.”
“Always check using this exact method.”
These may be useful heuristics.
They are not automatically standards.
If a criterion is presented as required, it should be grounded in the curriculum, task definition, assessment demand or clearly declared teaching purpose.
This is especially important for examination preparation.
Do not manufacture unofficial examiner rules.
Use current official sources when making specific claims about Singapore examination requirements.
Where no official rule exists, label the tutor method as a strategy, not a requirement.
Composite case: Mathematics success criteria that hide method selection
This case is fictional.
Ciara is learning simultaneous equations.
Her tutor gives a practice card:
1. Make one variable’s coefficients equal.
2. Eliminate.
3. Solve.
4. Substitute back.
5. Check.
Ciara performs well.
Later she encounters a pair of equations where substitution is much more efficient.
She still forces elimination because the visible criteria have become the task.
The tutor revises the criteria.
“Choose a valid efficient method.”
“Keep equivalent equations.”
“Solve both variables.”
“Check in the original equations.”
Now method selection re-enters the learner’s job.
The earlier checklist taught one procedure.
The later criteria define successful performance more broadly.
This is how criteria should evolve as expertise grows.
Feedback and criteria should meet
Feedback is most useful when the learner can locate it against the quality target.
“You need more detail” is weak.
“Your evidence is relevant, but the link to the claim is missing” is actionable because the criterion is visible.
“Check your answer” is weak.
“Your method is valid, but the final unit does not match the quantity asked for” is anchored to the standard.
Criteria make feedback interpretable.
Feedback makes criteria operational.
Then the learner should act.
Revise.
Retry.
Explain.
Check a fresh case.
The point is not to receive a judgement but to use the judgement to improve the next performance.
Criteria can support learner agency without handing over the standard
Learner voice does not mean the learner decides what counts as correct Mathematics or accurate Science.
But learners can participate in making standards usable.
After studying examples, ask them to propose criteria.
Compare their criteria with the curriculum or task standard.
Discuss which criterion is essential and which is stylistic preference.
This can deepen understanding of quality.
It can also reveal misconceptions.
A learner says, “A good explanation uses difficult words.”
The tutor now knows that vocabulary sophistication has been confused with explanatory quality.
Agency becomes calibration, not relativism.
Criteria should be tested against fresh work
A learner can recite criteria and still fail to use them.
After criteria are taught, watch whether they influence performance.
Does the learner choose more relevant evidence?
Do they check the unit?
Do they state the condition?
Do they revise the weak link?
Then test a fresh task with less visible support.
The receipt is changed performance.
Not agreement with the criteria.
Not highlighting the rubric.
Not correctly repeating “claim, evidence, link”.
A practical success-criteria protocol
State the learning goal.
Ask what successful performance needs to accomplish.
Separate enduring quality dimensions from exact answer content.
Decide how much of the criteria should be visible for the task purpose.
Use examples or contrasts if the criteria are too abstract to operate.
Ask the learner to apply the criteria during or after the task.
Fade criteria that are becoming cues.
Ask the learner to generate the criteria.
Return a fresh task under the target support condition.
Then report the conclusion honestly.
“Can use relevance and justification criteria independently on fresh inference questions.”
Not:
“Knows the rubric.”
Failure modes
The hidden-target failure. Learners are expected to improve without understanding what success means.
The model-answer checklist failure. Criteria reveal the exact content or method the learner is supposed to generate.
The compliance failure. Ticking boxes replaces judgement about quality.
The rubric overload failure. The criteria consume more attention than the task.
The strategy-as-standard failure. A tutor’s preferred heuristic is presented as an official requirement.
The permanent-scaffold failure. The checklist remains visible forever, so independent monitoring is never tested.
The premature-removal failure. Criteria are withdrawn before the learner has internalised the quality standard.
The criteria-drift failure. The wording stays the same while the actual judgement standard changes.
The one-size failure. The same criteria are used for teaching, diagnosis and verification without considering how the support changes the evidence.
The exam-myth failure. Unverified tutor rules are presented as examiner requirements.
Criteria transparency and three-learner tuition
Small groups create a useful opportunity.
Learners can compare criteria interpretations before comparing answers.
“Which criterion matters most here?”
“Did all three of us interpret ‘justify’ the same way?”
“Which answer meets the criterion most strongly and why?”
This can deepen quality judgement without turning peer work into answer copying.
But preserve independent first responses when they matter.
If learners see one strong answer before generating their own, later performance is partly supported by peer exposure.
The sequence can be:
Private attempt.
Criteria-based self-check.
Peer comparison.
Revision.
Fresh individual task later.
The criteria support discussion while the later receipt restores independence.
Criteria can become increasingly compressed
Novices may need a full explanation.
“Use evidence that directly supports the inference and explain the relationship.”
Later this can compress to:
“Evidence → link.”
Later still:
“Relevance?”
Eventually the learner may need no external cue.
Compression is useful because it reveals whether the underlying standard has become internal.
A short cue should not work by magic.
It works because the learner has built the fuller model behind it.
If the compressed cue stops producing the right monitoring behaviour, expand again temporarily.
Fading is evidence-responsive, not calendar-driven.
Separate process conditions from quality outcomes
Success criteria often fail because they mix two different things.
A process condition tells the learner what action to take while working.
A quality outcome tells the learner what the finished performance must accomplish.
Those can overlap, but they are not interchangeable.
“Underline the command word” is a process action.
“Answer the command actually asked” is a quality condition.
“Use the PEEL structure” is a process heuristic.
“Develop a relevant claim with supporting evidence and a clear link” is closer to the quality target.
“Draw a bar model” is a possible process.
“Represent the quantities and relationships correctly enough to solve the problem” is the outcome.
This distinction matters because a process can help a novice and later become optional.
If the tutor quietly turns the process into the standard, the learner may believe there is only one legitimate route to quality.
That is especially dangerous when several methods are valid.
A good tutor can therefore maintain two layers.
The first layer is temporary operating support:
“What should you do next?”
The second layer is the enduring standard:
“What must still be true when the support disappears?”
During early instruction, both layers can be visible.
During later verification, the process layer can fade while the quality layer remains.
Eventually, even the quality layer may need to be recalled by the learner rather than supplied externally.
This creates a clean progression from guided performance to independent judgement.
Criteria should change resolution as expertise changes
A criterion that helps a novice can become too coarse or too controlling later.
Early in algebra, “keep both sides balanced” may be the central quality cue.
Later, the learner must choose efficient transformations, preserve domain conditions, manage symbolic equivalence and decide when a result needs checking.
Early in writing, “give evidence and explain the link” can be useful.
Later, the learner needs to judge relevance, sufficiency, nuance, audience and the strength of competing interpretations.
The tutor should therefore review the resolution of the criteria.
Are they still naming the important decisions?
Have they become obvious?
Are they blocking a more advanced judgement?
Have they frozen one school-level scaffold after the learner is ready for a broader disciplinary standard?
This does not mean changing the standard arbitrarily.
It means making the visible criteria fit the learner’s present level of responsibility.
The enduring construct can remain stable while the external description becomes more sophisticated.
A mature learner should not carry a Primary-level checklist into every advanced task merely because it once helped.
Criteria are scaffolds for judgement as well as descriptions of quality.
They should develop with the judgement they are meant to support.
Parent transparency should explain the standard without creating an answer service
Parents reasonably want to know what the tutor is teaching and what improvement would look like.
Success criteria can make that communication clearer.
But there is a boundary.
If every home update contains the exact answer structure, exact vocabulary and exact steps the child should use next, the family can unintentionally become an extra prompting system.
A better parent explanation describes the current quality target and the evidence needed.
For example:
“Right now we are checking whether she can choose relevant textual evidence and explain the link independently.”
That is more useful than sending the exact sentence frame that the next task will test.
For Mathematics:
“We are checking whether he can select a valid method and preserve equivalent equations without being told which method to use.”
That keeps the educational job visible without giving the method decision away.
Parents can then support conditions around the work—time, materials, routine, encouragement—without becoming a second tutor delivering the missing cue.
Transparency should increase shared understanding.
It should not multiply assistance.
Programme consistency needs common standards, not identical scripts
A tuition programme may want tutors to use common success criteria so learners receive a coherent standard across groups.
That can be valuable.
But common criteria should protect the shared educational target, not force every tutor to use identical wording in every lesson.
One tutor may say, “What makes the evidence relevant?”
Another may say, “Show me why this quotation proves the claim.”
If both are orienting the learner toward the same quality dimension, variation in language can be legitimate.
Programme consistency matters most where different wording would actually change the standard.
If one tutor accepts any supporting detail and another requires evidence directly linked to the inference, learners are receiving different criteria.
If one tutor treats one preferred Mathematics method as mandatory while another accepts any valid efficient method, the programme has a standards problem, not merely a style difference.
A programme can therefore maintain a small common core:
the target capability;
the load-bearing quality dimensions;
examples of boundary cases;
what support is allowed at different task purposes;
and what later independent evidence is required.
Tutors retain professional language and adaptation around that core.
The result is coherence without scripting.
Evidence boundaries
AERO’s Explain learning objectives, published February 2024, is research-informed guidance supporting clear learning objectives and success criteria as part of explicit instruction and formative assessment.
AERO’s Formative assessment practice guide recommends clear measurable learning objectives and clearly explained success criteria. It is practice guidance, not evidence that any particular checklist format or fading sequence is optimal.
The National Student Support Accelerator’s Tutoring Session Structure is tutoring programme guidance that recommends framing objectives and aligned formative assessment. It supports coherence between goals, instruction and checking.
EEF’s Embedding Formative Assessment describes a programme whose key strategies include clarifying and sharing learning intentions and success criteria. The EEF page also reports the programme’s evidence context; that evidence should not be converted into a claim that any isolated success-criteria routine has the same effect.
EEF’s Feedback, published 24 August 2026, emphasises feedback relative to goals and helping learners understand success and next steps in the context of developing independent learners aged 16–19.
The transparency decisions in this Tutor Handbook article are professional principles. No claim is made that one criterion resolution, one rubric length or one fading sequence is universally correct.
The end state
A learner should not have to guess what quality means.
They should also not need the tutor’s checklist to perform every quality decision.
The tutor makes the standard visible.
The learner studies examples.
They learn the dimensions that matter.
They use the criteria to plan, monitor and revise.
Then the external structure becomes lighter.
The learner can say what good work requires.
They can recognise when their own work misses the standard.
They can make the correction without waiting for the tutor to point to the box.
And when the final task requires independent judgement, the criteria no longer need to sit beside the page performing that judgement for them.
That is the Success-Criteria Transparency Gate.
Make the target clear.
Keep the answer the learner’s.