A teacher gives a test, marks it carefully, records the scores and moves on.
Assessment happened. Teaching may not have changed.
The deeper purpose of assessment inside teaching is not simply to produce a number. It is to generate evidence about what learners know, what they can do, what support they still need, what the task actually measured, and what the teacher should do next.
Assessment for teaching works when evidence is deliberately collected, interpreted against the intended learning, and used to improve instruction, feedback, practice, grouping, support and future assessment—without pretending that one score can explain the whole learner.
This article continues the eduKate Sengkang How Teaching Works series after How Small-Group Teaching Works. It owns the teacher-side architecture of assessment as evidence for better teaching.
It does not replace subject-specific assessment guides, parent progress pages, learner-side marked-paper routes, or Tutor Handbook evidence/construct gates. Those pages retain their narrower jobs.
A route through assessment for teaching
Begin with assessment purpose, then examine what the task is meant to measure, evidence quality, sampling, formative assessment, summative evidence, interpretation, feedback, instructional response, and verification. Later sections cover Mathematics, Science, English, small groups, practice papers, rubrics, self-assessment, parents, AI, failure modes and evidence limits.
Assessment begins by naming the decision it should improve
Different assessments serve different decisions.
- diagnosis before teaching;
- checking understanding during teaching;
- deciding whether to move on;
- deciding who needs additional support;
- identifying misconceptions;
- checking retrieval after delay;
- testing transfer to a new problem;
- judging independent performance;
- summarising attainment;
- preparing for an examination;
- evaluating whether an intervention changed performance.
A useful assessment should produce evidence relevant to the decision being made.
One task cannot do every assessment job equally well
A short exit question can guide tomorrow’s lesson. It cannot establish broad mastery of an entire subject.
A full examination can sample many topics but may give weaker evidence about the exact reasoning behind one mistake.
Use the smallest assessment that can support the decision, and use broader evidence for broader claims.
Assessment should be designed before the lesson needs the answer
Teachers can plan key evidence points in advance: prerequisite check, hinge question, first independent example, exit task, delayed return.
Preplanning reduces the temptation to ask convenient questions that generate answers but do not help the teacher decide what to do.
Define the construct: what capability is the assessment supposed to reveal?
The construct is the capability or knowledge the teacher wants evidence about.
If the target is fraction magnitude, dense reading language can contaminate the measurement. If the target is written argument, an oral answer cannot fully replace written evidence. If the target is method selection, a worksheet title that names the method changes the construct.
Assessment quality improves when the task conditions preserve the target and reduce irrelevant barriers.
Construct underrepresentation is a common classroom problem
A narrow task samples too little of the target.
One correct inference does not establish broad comprehension. One successful equation does not establish algebraic flexibility. One well-written paragraph does not prove sustained composition control.
Broader claims need broader samples.
Construct-irrelevant difficulty can distort evidence
A learner may know the Science but misread the complex vocabulary of the question. A learner may understand the Mathematics but lose marks because the graph is visually inaccessible.
Where those features are not part of the intended capability, adapt access or use a second task to separate them.
Supported performance and independent performance are different constructs
A correct answer after a prompt is useful evidence of learning under support.
It is not identical to a correct answer produced without that prompt.
Assessment should record which state it observed.
Good assessment evidence is interpretable, not merely scorable
A score is easy to record. A useful assessment tells the teacher something about the next instructional decision.
The evidence becomes stronger when the teacher can see relevant reasoning, error location, confidence, support conditions, method selection or transfer behaviour.
Correct answers can hide weak reasoning
Guessing, copied methods, accidental cancellation of errors and peer cues can produce correct outcomes.
Use justification or a fresh parallel task when the reasoning matters.
Wrong answers can hide strong partial knowledge
A learner may select the correct method, build the right representation and make one arithmetic slip.
Do not reteach the whole topic when the error mechanism is local.
Blank answers are ambiguous evidence
A blank can mean missing knowledge, time shortage, task avoidance, misunderstanding, low confidence or strategic skipping.
Follow up before treating omission as a specific knowledge diagnosis.
Confidence can enrich evidence without replacing performance
Ask how confident the learner is and compare that with the response.
High-confidence wrong answers may need misconception repair. Correct low-confidence answers may need retrieval and calibration.
Assessment samples performance; it does not contain the whole capability
Every assessment samples some tasks from a larger domain.
Teachers should therefore ask whether the sample is wide enough, varied enough and representative enough for the claim.
Sample across important content, not only convenient questions
Question banks can overrepresent familiar, easily marked tasks.
Include the important knowledge structures, even when they require richer responses.
Sample across representations
If a concept can appear as prose, graph, diagram, equation or table, use more than one form when transfer matters.
Sample across support conditions
Early assessment can be supported. Later assessment should remove the support relevant to the target claim.
This is how assessment connects with scaffolding and fading.
Sample across time
Immediate success and delayed success answer different questions.
Return later when durable learning is the claim.
Sample across variation
Near-copy success is useful but should not be confused with transfer.
Use changed wording, representation and context when portability matters.
Formative assessment is evidence used to move learning forward
Formative assessment is defined by use, not by the shape of the activity.
A quiz can be formative if the teacher uses the evidence to adapt teaching. A conversation can be formative. A draft can be formative. A practice paper can be formative.
AERO’s current formative assessment guidance and Monitor Progress guide emphasise low-stakes evidence used to identify what learners understand and to target instruction, guidance or feedback.
A formative question needs a planned response
Before asking a hinge question, know what different answer patterns might trigger.
If all learners move to the same next task regardless of answer, the question may still provide practice, but its formative function is weak.
Formative assessment should not become constant interruption
Teachers do not need to assess every sentence of learning.
Check at consequential points: prerequisite, new distinction, first application, before support fades, before progression, after feedback, after delay.
Formative assessment should include what learners can do correctly
Assessment should reveal readiness as well as weakness.
Strong evidence may justify extension, reduced support or longer spacing.
The Embedding Formative Assessment evidence is useful but specific
EEF reports that the tested secondary-school Embedding Formative Assessment programme produced an average gain equivalent to two months of additional progress in Attainment 8, with a very high security rating.
That is evidence about a particular professional-development programme and implementation context, not proof that every classroom technique labelled formative assessment has the same effect.
Summative assessment can still inform teaching
An end-of-unit test or examination may have a summative purpose and still generate useful diagnostic evidence afterward.
Teachers can analyse error classes, omissions, timing, transfer and content patterns to improve future teaching.
Do not overdiagnose from a total score
62% can arise from many different knowledge profiles.
Inspect where marks were lost and what the loss means before prescribing a route.
Subscores can also mislead when based on too few items
A “weak algebra” label based on two questions may be unstable.
Use enough evidence before broad claims.
Practice papers should be read as integrated evidence
They combine content, retrieval, selection, time, reading, checking and transfer.
Classify errors before deciding what training follows.
Interpret assessment at the level of the mechanism
Useful error classes include:
- missing knowledge;
- weak retrieval;
- misconception;
- representation failure;
- method-selection failure;
- execution error;
- language or task-interpretation barrier;
- checking failure;
- transfer failure;
- timing failure;
- support dependence.
The next teaching move depends on which mechanism best explains the pattern.
Keep observation, interpretation and decision separate
Observation: “The learner omitted the second term when expanding the bracket.”
Interpretation: “A distributive-property misconception is possible.”
Decision: “Use a contrast item and fresh check before broader reteaching.”
This language keeps diagnosis revisable.
Look for patterns across tasks
One error can be noise. Recurring error under related conditions is stronger evidence.
Use multiple data points before large route changes.
Beware correlated evidence
Five near-identical worksheet successes are not five independent confirmations of broad mastery.
Vary task form and timing before confidence grows.
Beware difficulty drift
Scores can change because question difficulty changes.
Compare like with like when tracking progress, or record how the task changed.
Beware support drift
Performance may improve because more prompts, tools or peer help are available.
That can be real learning opportunity, but it changes the measurement.
Assessment should lead into feedback or teaching
Evidence without response becomes record keeping.
Use assessment to decide whether the learner needs feedback, explanation, guided practice, retrieval, misconception repair, transfer work or extension.
See How Feedback Works in Teaching.
Assessment should give learners something to do
Repair, explain, compare, reattempt, retrieve, self-check or solve a fresh case.
Do not make “receive mark” the end of the learning loop.
Assessment should change instruction when the evidence warrants it
If many misunderstand, pause and fix. If some are unsure, adapt support. If most understand, extend.
This connects assessment directly to How Adaptive Teaching Works.
Not every assessment result needs action
A stable, low-consequence slip may not justify a new teaching route.
Prioritise what most affects the learning goal.
Assessments can also tell the teacher to reduce support
Strong independent evidence should change teaching too.
Fade prompts, lengthen spacing intervals, increase variation or move to transfer.
Diagnostic assessment before teaching should be short and discriminating
Do not pre-test an entire unit if three well-chosen items can reveal the prerequisite states that matter.
Use the smallest set that can distinguish likely routes.
Hinge questions should sit where the lesson can genuinely change
A hinge question before independent practice can show whether the class is ready for reduced support.
Plan the response to common answer patterns before asking.
Exit tasks should inform the next lesson
“What did you learn?” may support reflection, but a task-specific item often gives stronger instructional evidence.
Read the responses before deciding the next lesson’s opening.
Rubrics should clarify quality without shrinking it into compliance
A rubric can make criteria visible. It can also encourage box-ticking if every quality dimension is reduced to surface features.
Use rubrics to support judgement, not replace it.
Mark schemes are evidence rules, not magic keyword lists
Teach learners why an answer earns credit: relevant knowledge, valid reasoning, evidence, method, accuracy or communication.
Do not convert mark schemes into superstition about isolated words.
Self-assessment should be calibrated against external evidence
Ask learners to judge one aspect of their work, then compare with teacher or criterion evidence.
The goal is improving self-monitoring, not asking learners to award their own final grade.
Peer assessment needs a narrow, teachable job
“Make your partner’s work better” is too broad.
“Identify the claim, locate the evidence and check whether the link is explained” is more interpretable.
The original learner should retain authorship of the repair.
Worked case: Mathematics assessment separates method selection from arithmetic
Question: solve 4(x – 2) = 24.
Learner A writes x – 2 = 6, then x = 8: correct method and execution.
Learner B expands to 4x – 8 = 24 then writes 4x = 16: arithmetic/transposition error.
Learner C writes 4x – 2 = 24: distributive-property issue.
Same item, different instructional implications.
Worked case: Science assessment separates data reading from causal explanation
A graph shows temperature changing over time under two conditions.
First ask for the numerical change. Then ask what pattern the data show. Then ask for a mechanism. Then ask what conclusion would be too strong.
Each question samples a different layer of scientific reasoning.
Worked case: English assessment separates evidence selection from expression
A learner gives a weak written inference answer.
Ask orally what the passage suggests and which detail supports it.
If the oral reasoning is strong, writing expression may be the barrier. If the interpretation remains unsupported, the reading reasoning itself needs repair.
Small-group assessment can be precise without becoming constant testing
Use short individual entry checks, observe guided practice, sample one independent item and record support state.
The group should still have time to learn, discuss and practise.
Progress should be compared against meaningful baselines
Compare like conditions when possible: similar difficulty, similar support, similar timing and similar construct.
If conditions change, document the change rather than pretending the scores are directly equivalent.
Parents need evidence translated into learning decisions
“Scored 72%” is less useful than “retrieval is stable, but method selection breaks on mixed questions; next block targets discrimination.”
Report what was observed, what it likely means, and what will be checked next.
Do not claim causation from ordinary progress data
A score increase after tuition may reflect tuition, school teaching, maturation, practice, test difficulty or several factors together.
Report improvement accurately without claiming more causal certainty than the evidence permits.
AI can generate assessments quickly and generate bad evidence quickly
AI can create question variants, rubrics, distractors and practice sets.
Teachers should verify answer correctness, ambiguity, difficulty, curriculum alignment and whether the question actually measures the intended construct.
Do not allow AI to infer permanent learner traits from a few responses.
Digital adaptive tests can change difficulty and therefore change comparability
If one learner receives harder items because earlier answers were correct, raw percent scores may not be directly comparable.
Understand the platform’s scoring and evidence model before using the numbers for high-stakes instructional decisions.
Assessment data should be stored only as long as it serves an educational purpose
Keep task-relevant records inside appropriate authorised systems. Avoid speculative personal labels.
Update records when the learner changes.
Assessment systems need an evidence hierarchy
- one supported example: narrow evidence;
- fresh independent example: stronger evidence;
- several varied examples: broader evidence;
- delayed retrieval: durability evidence;
- mixed selection: discrimination evidence;
- changed context: transfer evidence;
- integrated realistic task: performance evidence.
Use the evidence level that matches the claim.
A practical teacher sequence for assessment
- Name the decision the assessment should improve.
- Define the construct.
- Choose a task that samples the construct without avoidable contamination.
- Decide what support conditions are appropriate.
- Collect enough evidence for the size of the claim.
- Inspect reasoning, not only score.
- Separate observation from interpretation.
- Classify the likely mechanism behind errors.
- Choose feedback, teaching, practice, support or extension accordingly.
- Give the learner an action.
- Use a fresh task to verify the repair.
- Return after delay or variation where durability matters.
- Update the learner model and next teaching plan.
Common assessment-for-teaching failure modes
- collecting scores without a decision plan;
- using one task to make broad mastery claims;
- ignoring support conditions;
- confusing task difficulty with learner ability;
- treating every wrong answer as missing knowledge;
- treating blanks as one specific diagnosis;
- using chapter-labelled questions to claim method selection;
- grading group work as individual understanding;
- averaging unlike evidence into one score;
- comparing scores from very different difficulty levels as if equivalent;
- assessing constantly and teaching too little;
- returning marks without response time;
- using rubrics as compliance checklists;
- using AI-generated items without validation;
- keeping old learner labels after evidence changes.
Assessment should end by improving the next evidence cycle
After instruction changes, assess again in a way that tests the targeted capability.
If the learner succeeds only with the same cue, the repair may be incomplete. If success survives fresh conditions, confidence grows.
Evidence, interpretation and limits
AERO’s Monitor Progress guide, last updated 14 May 2026, emphasises checking what students understand and can apply, identifying gaps, and adjusting instruction through additional teaching, guidance or feedback. Its July 2026 formative-assessment resources frame low-stakes evidence as a tool for targeting instruction.
EEF’s Embedding Formative Assessment effectiveness trial involved 140 schools and approximately 25,000 pupils and reported an average gain equivalent to two additional months of Attainment 8 progress, with a very high security rating. That result is programme-specific and does not justify treating every formative-assessment routine as equally effective.
The assessment architecture, evidence ladder and worked cases in this article are editorial teaching designs. They have not been evaluated together as one intervention. Assessment validity and reliability also depend on context, task design, administration and interpretation.
Sources for this edition were reviewed on 22 September 2026.
The assessment-for-teaching standard: evidence should change what happens next
Assessment is not complete when the score is recorded.
The teacher should know what the task measured, what the evidence permits, what remains uncertain, and which next move is justified.
Assessment serves teaching when it turns learner performance into a more accurate next instructional decision.
Continue through How Teaching Works, revisit How Adaptive Teaching Works, or use How Feedback Works in Teaching for the response loop.