The Tutor Handbook · Volume 0109 · Series ID THB-0109
A tutor writes five criteria on the board before a task: identify the relevant concept, choose the correct method, show the working, check the units, explain the final answer. The learner completes the task beautifully. The page is accurate, organised and complete. The tutor is pleased.
Then the criteria disappear.
On the next question the learner stalls at the first decision. The problem was not that the criteria were wrong. The problem was that some of them were doing more than making quality visible. They were naming the route. What looked like a transparent standard had become a partial answer key.
The direct answer is that tutors should disclose criteria according to the job of the task. Criteria should make the destination, constraints and standard of quality visible. They should not quietly perform the learner operation that the task is meant to test. If the target is method selection, a criterion that names the method is too revealing. If the target is evidence use, a criterion such as “support the claim with relevant evidence and explain the connection” may clarify the standard without telling the learner which sentence to choose. If the target is a routine that the learner is still acquiring, a more explicit checklist may be justified during teaching—but the tutor should later remove or compress it before claiming independent performance.
Why success criteria can both help and contaminate
Learning is easier to regulate when the learner knows what good work is supposed to achieve. Vague standards force students to guess what the adult is looking for. A learner told only to “write better” receives little usable information. A learner told that a strong explanation must state the relevant observation, connect it to an appropriate concept and make the causal relationship explicit has a clearer quality target.
But transparency has a boundary. The more a criterion names the exact sequence of moves, the more it can become a scaffold. Scaffolds are not bad. They are often necessary. The measurement problem begins when supported performance is described as though the support did nothing. A criterion can be legitimate teaching support and still be too informative for a later independence claim.
This article therefore owns a narrow tutoring decision. It does not replace the general explanation of formative assessment at How Formative Assessment Works, which already covers learning goals, success criteria and evidence use. The Criteria Disclosure Gate asks something more operational: how much of the standard should be visible at this moment, in this task, when some forms of visibility may also reveal the route?
Destination information and route information
A useful distinction is between destination information and route information. Destination information tells the learner what successful work must accomplish. Route information tells the learner what to do next, which method to select or which features to attend to in a particular order.
“Your answer must compare both sources and justify which is more reliable” is mainly destination information. “First identify the author, then check the date, then compare purpose, then choose Source B” is largely route information. “Use an appropriate method and justify your choice” describes a standard. “Use substitution” supplies a decision. “Explain the relationship between the observation and the scientific concept” describes the job. “Write that particles move faster, collide more often and therefore…” may be a worked answer.
The distinction is not always clean. A novice often needs route information while learning. An expert may need only destination information. The same checklist can be helpful on Tuesday and too revealing on Friday if the educational claim has changed from guided acquisition to independent selection. That is why disclosure is a gate rather than a permanent formatting rule.
Four questions before showing the criteria
- What is the target operation? What must the learner eventually choose, retrieve, organise, explain or check without the tutor?
- What must be transparent for fairness? Which requirements would be unreasonable to hide because they define the task or expected quality?
- What part of the criteria carries the target? Does any line tell the learner the very method, cue, structure or sequence being assessed?
- What claim will follow? Are we teaching with support, rehearsing with support, or checking what the learner can initiate independently?
These questions stop the tutor from treating “success criteria” as a single educational object. Criteria can be a teaching scaffold, a quality reference, a self-checking tool, a marking explanation or an assessment condition. Those jobs overlap, but they are not identical.
The five disclosure levels
A practical way to think about disclosure is to use five levels. These are not validated psychometric levels and should not be scored. They are a planning language for tutors.
Level 1: Full model. The learner sees a complete example, annotated criteria and the route that produced the answer. This is teaching. It may be exactly right for first contact with a demanding structure.
Level 2: Process checklist. The learner sees the major stages or decisions but must execute them. This is guided practice. It can build a routine, but the tutor should note which decisions remain externally named.
Level 3: Quality criteria. The learner sees what the finished response must accomplish but not the particular method or sequence. This can support self-monitoring while preserving more route selection.
Level 4: Compressed standard. The learner receives only a short quality reminder, such as “complete, justified and checked” or a subject-specific equivalent. The learner must reconstruct the underlying criteria.
Level 5: Internalised standard. The visible criteria are absent. The learner must identify what quality requires and inspect the work accordingly. This is not automatically superior. It is appropriate only when the task genuinely requires the learner to carry the standard themselves and legitimate access supports remain available.
Composite case: Alicia and the Mathematics checklist
This is a fictional composite case. Alicia is learning to solve unfamiliar percentage-change problems. Her tutor gives a checklist: identify original value, identify new value, find the difference, divide by original value, multiply by 100, state increase or decrease. With the checklist visible, Alicia completes eight questions accurately.
The tutor could reasonably say that Alicia can execute the taught process with a checklist. The tutor cannot yet say that Alicia can recognise and organise a percentage-change problem independently. The checklist performs several high-value decisions: it names the denominator, orders the operations and reminds her how to classify the result.
The next stage therefore compresses the criteria rather than removing all support at once. Alicia receives: “Identify the reference quantity, compare the values, express the change relative to the correct base, label the direction.” She has to choose the arithmetic. Later the visible criteria disappear and a mixed set includes ordinary percentages, reverse percentages and percentage change. The tutor now observes whether Alicia can decide what kind of percentage problem she is facing before executing.
The checklist was not a mistake. It was a scaffold. The mistake would have been calling checklist-guided execution independent method selection.
Composite case: Beatrice and the Science keywords
Fictional composite case. Beatrice is asked to explain why dissolving happens faster under one condition. The tutor has a criterion sheet that lists “greater particle movement”, “more frequent contact”, “faster dissolving”. Beatrice writes a polished explanation using all three phrases. On a changed question about a different variable, she cannot decide which mechanism matters.
The criterion sheet has crossed into content provision. It did not merely say what a causal explanation should do; it supplied the mechanism-specific ideas that the learner was supposed to retrieve and connect. During initial teaching that may be useful. During a diagnostic or transfer task it is too revealing.
The tutor replaces the keyword list with a generic quality standard: “Name the relevant observation or change, identify the scientific relationship that explains it, and connect cause to effect precisely.” Now Beatrice must select the concept. If she cannot, the tutor has cleaner evidence about what is missing. After teaching, the tutor can reintroduce a checklist for revision if it serves self-monitoring, but should later check performance without mechanism-specific prompts.
Composite case: Ciara and the English writing frame
Fictional composite case. Ciara is working on situational writing. Her planning sheet names audience, purpose, required content points, tone, opening move, paragraph sequence and closing move. Her response is complete. The tutor then gives her a new task without the frame. She misses one content requirement and uses a tone that does not fit the relationship.
The frame has been doing two jobs. It made the task requirements visible, which is useful. It also performed the planning sequence. The tutor now has to decide which part should remain visible. If the formal task itself supplies explicit content requirements, hiding those would be artificial. But the tutor can remove the paragraph-order prompts and ask Ciara to generate her own plan from purpose, audience and task conditions.
Later, the tutor might let Ciara inspect a short criteria card only after the first draft. That changes the job from “follow this plan” to “audit your own work against a standard”. The same criteria become less route-giving because their timing has changed.
Timing changes what a criterion does
Criteria shown before a task can orient attention and reduce ambiguity. Criteria shown during a task can scaffold monitoring. Criteria shown after a first attempt can support self-assessment without shaping the initial route as strongly. Criteria shown after feedback can help the learner reconstruct why the feedback matters. Criteria revisited after a delay can reveal whether the learner has internalised the standard.
This means tutors do not have to choose between total transparency and total concealment. They can change timing. A learner may attempt first from memory, then open a criteria card to audit the answer, then close the card and repair independently. That routine produces more information than either permanent criteria visibility or a sudden blank page.
Criteria for quality are different from answer cues
A good criterion often names a property of strong work rather than the content of this particular answer. “Use evidence that directly supports the claim” is a quality property. “Use paragraph three because it mentions the storm” is an answer cue. “Justify why the selected method fits the structure” is a quality property. “Use completing the square” is a method cue. “Connect observation, concept and causal relationship” is a quality property. “Mention heat transfer by conduction” is content provision when conduction is what the learner is meant to identify.
There are exceptions. If the lesson objective is to practise executing completing the square after method selection has already been taught, naming the method is not contamination; it defines the practice set. If the lesson objective is to recognise when completing the square is useful, naming it destroys the recognition demand. The wording is the same. The educational job changes.
Criteria can create a hidden dependency
A criteria sheet can become so familiar that the learner stops constructing a mental model of quality. Instead, the learner performs a visual scan: box one done, box two done, box three done. This can be useful during early routine building. Over time, however, the external checklist may become the only place where the standard exists.
The warning signs are practical. The learner asks for the checklist before reading the task. They cannot explain why a criterion matters. They complete the boxes in order even when the task does not require that order. They become anxious when the criterion wording changes. They can identify missing boxes on a page but cannot judge the quality of an unfamiliar response. These signs do not mean checklists are harmful. They mean the handover has not finished.
A fade can therefore target the representation of the standard. Full checklist becomes short labels. Short labels become learner-generated questions. Learner-generated questions become silent self-monitoring. The tutor occasionally asks the learner to make the criteria explicit again, not because the criteria must always remain visible but because explaining them reveals whether the learner understands the quality standard rather than memorising the card.
Do not hide requirements in the name of independence
The opposite error is equally serious. Some tutors try to test independence by withholding information that a fair task should make clear. If a question has a required output form, marking condition or audience, the learner should not have to guess it unless interpreting that condition is itself the target. Independence does not mean solving an adult’s private puzzle about what counts as success.
Similarly, legitimate access arrangements should not be withdrawn to make a task feel more authentic. If a learner requires an access support that does not perform the target capability, the tutor should preserve it. The disclosure gate concerns instructional information that may carry the target operation. It is not a licence to remove accessibility, change official assessment conditions or deny reasonable supports.
The criteria audit
Before giving a criteria sheet, the tutor can audit each line with four labels:
- Requirement: defines what the task explicitly asks for.
- Quality: describes what strong work is like.
- Route: tells the learner a procedure or sequence.
- Answer content: supplies a concept, method, evidence source or conclusion that may be the thing the learner is meant to generate.
Requirements and quality criteria can often remain visible. Route information may be appropriate during teaching and need fading later. Answer content should be treated very cautiously if the task is supposed to reveal independent retrieval, selection or reasoning. A single criterion can contain more than one category, which is why the audit is about function rather than wording alone.
Three learners can need different disclosure without different standards
In a three-student tutorial, fairness does not require identical scaffolds. Alicia may use a full planning frame while she is acquiring a routine. Beatrice may use compressed criteria because she can initiate the route but misses a quality condition. Ciara may attempt without visible criteria and then use them for post-attempt auditing. All three can still be working towards the same underlying standard.
The tutor should record the support condition because the same final score does not mean the same thing under different disclosure levels. That does not turn support level into a human ranking. It simply preserves evidence provenance. The goal is to move each learner towards carrying more of the relevant standard internally where that is educationally appropriate.
When criteria should become more explicit, not less
Fading is not always the next move. If a learner repeatedly misunderstands what quality requires, more explicit criteria may be necessary. If feedback conversations keep circling around hidden expectations, the tutor should make the standard visible. If two markers would reasonably disagree about whether a response is sufficient because the standard is vague, learners should not be expected to infer an invisible rule.
The disclosure gate is therefore bidirectional. Sometimes the tutor reduces route information to test ownership. Sometimes the tutor increases quality information to remove needless ambiguity. The decision depends on what uncertainty is currently limiting learning.
Research: clarity of goals and criteria
The Australian Education Research Organisation’s Explain learning objectives practice guide recommends making learning objectives and success criteria clear and understandable to students, connecting them to prior knowledge and using them during lessons. AERO’s current video guidance likewise shows teachers referring learners back to objectives and criteria. This is research-informed school guidance; it does not test the exact disclosure levels proposed in this article or establish a universal amount of information that private tutors should reveal.
AERO’s Scaffold practice guidance is also relevant because criteria can function as scaffolds. It describes planned and contingent support and the need to adjust or fade assistance as learners become more capable. Its Monitor progress guidance reinforces the need to use evidence of student understanding to decide what to teach next. Again, these principles travel more safely than any claim that a particular checklist format will cause a particular outcome.
Research: tutoring and implementation
Stanford’s National Student Support Accelerator places structured instructional materials, instructional practices and formative assessment inside its Tutoring Quality Standards, while carefully distinguishing research-based, research-informed and emergent standards. That distinction matters here. Clear criteria and structured supports are research-informed components of strong instruction, but the micro-decision about exactly when a criterion becomes an answer cue is a professional judgement that requires attention to the target task.
The Education Endowment Foundation’s Effective tutoring summary reports positive average tutoring effects but stresses implementation and monitoring. Its implementation guidance treats successful implementation as an ongoing process of making an approach work in ordinary practice. Neither source validates this article’s disclosure ladder. They support the broader principle that materials and routines need to be enacted, monitored and adapted rather than treated as intrinsically effective objects.
The learner should eventually be able to generate the standard
A powerful test is to ask the learner, before seeing the tutor’s criteria, “What would a strong answer need to do?” The answer reveals whether the standard has moved inside. In Mathematics, the learner may say that a complete solution needs a justified method, accurate execution and a check that the answer satisfies the original conditions. In English, the learner may name audience, purpose, evidence and clarity. In Science, the learner may distinguish observation from explanation and identify the need for a causal bridge.
The learner-generated standard will not always match the official or instructional standard perfectly. That mismatch is useful. The tutor can compare the two, repair missing conditions and ask why each criterion matters. This turns criteria from a compliance sheet into a model of quality.
The changed-condition check
After a learner succeeds with criteria, change one condition. Remove one route-giving line. Reword the criteria. Ask the learner to generate the criteria. Show the criteria only after the first attempt. Use a new task where the same quality standard applies but the surface is different. The aim is not to trick the learner. It is to discover which part of the performance belonged to the learner and which part still belonged to the external representation of the standard.
If performance collapses, restore the smallest useful support. A failed fade is information. It may show that the learner knows the steps but cannot yet initiate them, understands the criteria but cannot retrieve them, or can judge quality after the fact but cannot plan towards it. Those are different teaching problems.
A parent-facing explanation
Parents often see a checklist and reasonably assume that it is simply good organisation. It may be. A tutor can explain the deeper issue without jargon: “This checklist is helping your child practise the routine. I am not yet using the checklist-guided work as proof that she can plan the whole response alone. Over the next few sessions I will shorten the prompts and check whether she can reconstruct the standard herself.”
That explanation prevents two unhelpful reactions. One is to celebrate a supported page as complete mastery. The other is to remove support abruptly because “she should know it by now”. The better route is planned transfer of responsibility, with evidence at each reduction.
The tutor’s own criteria can become too complicated
Criteria can fail through excess as well as disclosure. A page with eighteen boxes may contain accurate advice but impose a monitoring load larger than the task. Learners begin managing the checklist instead of the subject. The tutor may then misread checklist fatigue as weak knowledge. This is especially likely when criteria from several teaching episodes accumulate without retirement.
Compression is therefore part of expertise. As routines become stable, several criteria can be bundled into one higher-order question. “Have I answered the actual question with justified evidence?” can eventually replace five separate comprehension prompts. “Does every transformation preserve the relationship and answer the requested quantity?” can replace a long Mathematics audit. Compression should happen only when the learner understands what the short phrase contains.
Criteria are not a fourth tuition mode
The Tutor Classification Model remains a set of functions, not ranks. A Class 1 Explainer may use fuller criteria to make quality visible. A Class 2 Drill Builder may use a short process checklist while stabilising execution. A Class 3 Diagnostic Tutor may deliberately reduce route-giving criteria to see which decision the learner can make. A Class 5 Performance Coach may align visible criteria with realistic performance conditions. The function changes what the criteria are for.
Repair, Alignment and Frontier remain the three tuition modes. The same disclosure logic can operate inside each. Repair may temporarily need explicit scaffolds. Alignment may need clear standards so school and tuition expectations do not diverge. Frontier work may use broad quality constraints while preserving genuinely difficult decisions. There is no need to rename support as a new mode.
A compact Criteria Disclosure Gate
- State the target capability.
- Separate task requirements from quality criteria.
- Mark any line that names a method, sequence, cue or answer content.
- Decide whether the present job is teaching, guided practice, self-checking or independent evidence.
- Keep legitimate access supports.
- Show only as much route information as the present job justifies.
- Record the disclosure condition when interpreting performance.
- Reduce or retime criteria deliberately, not theatrically.
- Use a delayed or changed task to test whether the quality standard has become learner-owned.
This is not a score and should not become a bureaucracy. It is a way of keeping one deceptively simple question alive: did the criteria clarify what good work is, or did they quietly do part of the work?
The final return
A learner deserves to know what quality means. They also deserve the chance to become the person who can recognise that quality without carrying a laminated answer route forever. Good criteria make the standard more visible. Good tutoring then transfers control of that standard gradually and checks what survives.
The Criteria Disclosure Gate protects both sides of the problem. It prevents hidden expectations from becoming an unfair guessing game, and it prevents well-meant transparency from turning method selection, planning, retrieval or reasoning into checklist-following. The tutor can teach explicitly and still preserve honest evidence. The key is to know which information defines success and which information performs the learner’s next decision.
Make the destination visible. Teach the route when teaching requires it. Then, when the educational claim shifts to independence, stop confusing a map held by the tutor with a route the learner can now navigate alone.
Sources and evidence boundary
- Australian Education Research Organisation — Explain learning objectives: current research-informed practice guidance on clear learning objectives and success criteria.
- AERO — Scaffold practice: guidance on planned and contingent scaffolds and their adjustment as students become more capable.
- AERO — Monitor progress: guidance on using checks for understanding to adapt teaching and feedback.
- Stanford National Student Support Accelerator — Tutoring Quality Standards: tutoring-specific evidence framework covering instructional practices, materials and formative assessment while distinguishing evidence strength.
- Education Endowment Foundation — Effective tutoring: tutoring evidence summary with implementation cautions.
- Education Endowment Foundation — Implementation guidance: school implementation framework relevant to monitoring how routines work in practice.
The disclosure levels and gate in this article are an eduKate tutoring framework, not a validated scale and not an official MOE or SEAB assessment protocol. The cited guidance is largely school-based and international. It supports the underlying instructional principles but does not establish causal effects for a Singapore three-student tuition setting. All learner cases are fictional composites.