Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

How Assessment Evidence Works | What a Test Can Tell Us—and What It Cannot

Direct Answer: Assessment evidence works by giving us a structured observation of performance under defined conditions. The evidence becomes useful only when we ask what the task actually sampled, how consistently it was measured, what conditions shaped the result and how large a conclusion the evidence can support. A score is not the learner. It is one piece of evidence about one performance.

The simplest definition of assessment evidence

Assessment evidence is information produced by a task or observation that helps us make a bounded judgement about learning or performance.

In one line: Every score answers a question—but not necessarily the question we are tempted to ask.

The core mechanism

CLAIM → TASK → PERFORMANCE → OBSERVATION / MARKING → SCORE OR DESCRIPTION → INTERPRETATION → DECISION → CHECK AGAINST OTHER EVIDENCE

The first step is often forgotten: what claim are we trying to support? “The learner scored 72%” is a result. “The learner has mastered Primary 6 Science” is a much larger claim. Whether the result supports that claim depends on what the paper sampled, how it was constructed, how it was marked and the conditions under which the learner performed.

The professional Standards for Educational and Psychological Testing, developed by AERA, APA and NCME, place validity at the centre of test interpretation. In practical parent language: we should not ask a measurement to prove more than it was built to measure.

The eight questions to ask about assessment evidence

1. What was the assessment trying to measure?

A vocabulary quiz, a timed Mathematics paper, an oral examination and a Science practical sample different performances. Before interpreting a result, identify the target construct or capability.

Bolt 03 — What Exactly Did This Test Measure? exists because accurate marking does not guarantee accurate interpretation.

2. What content and skills were sampled?

No school paper contains every possible question from a subject. Tests sample. A paper with too few questions can exaggerate the importance of a narrow slice of content or miss parts of the capability entirely.

Bolt Measurement Note 13 explains why a small sample can misrepresent a broad skill.

3. Were the conditions relevant?

Timing, support, calculators, prompts, familiar formats, noise and fatigue can all alter the performance that appears. If the real future performance will be timed and independent, an untimed heavily supported task cannot fully establish readiness for it.

That does not make the supported task useless. It simply means it answers a different question.

4. How reliable is the result?

Educational performance contains variability. Different items, different days and different markers can produce somewhat different results. Reliability concerns the consistency and precision of the measurement, not whether the student “deserves” the score.

Bolt Measurement Note 11 — A Test Score Is an Estimate, Not an Exact Point is the parent-facing version of this measurement principle.

5. Are two scores actually comparable?

A rise from 60% to 75% looks like improvement, but the conclusion depends on difficulty, coverage, support, timing and marking. If those changed, the raw percentages may not represent the same measurement scale.

Bolt Measurement Note 02 protects against false progress claims built from incomparable tasks.

6. What does the score hide?

Totals compress. Two students can receive the same score while missing completely different operations. A class average can rise while some students fall behind. The same average can hide very different spreads. A perfect score can hit a ceiling and stop distinguishing further growth.

Assessment evidence becomes more diagnostic when we decompose the total back into patterns, items, processes and conditions.

7. What other evidence agrees or disagrees?

A test is one observer. Teacher judgement, learner prediction, previous work, delayed performance and fresh transfer tasks can support or challenge its interpretation. Agreement across different relevant sources can strengthen a conclusion; disagreement is a signal to investigate.

This is where assessment joins How Learning Calibration Works.

8. What decision is the evidence being used for?

The amount and quality of evidence needed should match the consequence of the decision. Choosing tomorrow’s practice problem requires less evidence than making a high-stakes judgement about long-term capability.

Higher-stakes claims deserve stronger, more relevant and more repeated evidence.

Formative and summative evidence

Formative assessment gathers evidence to improve learning while there is still time to act. The Education Endowment Foundation describes formative assessment as using activities that provide information about pupils’ understanding without contributing to formal grades. The evidence is valuable because it changes teaching or learning next.

Summative assessment summarises performance at a point in time, often for reporting, certification or progression. It can still inform future learning, but its primary job may be different.

The same task can sometimes serve both purposes, but we should know which job we are asking it to do.

What assessment evidence is not

  • A score is not a human value.
  • A correct answer is not proof of independent mastery if support was present.
  • A low score is not automatically evidence of low ability.
  • A high score is not proof that every relevant skill is strong.
  • A percentile is not the same as absolute achievement.
  • An average is not every learner.

A parent’s evidence table

EvidenceWhat it can help tell usWhat it cannot establish alone
One school testPerformance on sampled content under those conditionsFull capability across the subject
Repeated similar testsPattern under comparable tasksTransfer to substantially different tasks
Teacher observationProcess, misconceptions, response to instructionEverything the learner can do outside that environment
HomeworkPractice performance and work habitsIndependent performance unless support is known
Timed full paperIntegrated performance under more realistic exam constraintsHuman worth or every future performance

How assessment evidence works across subjects

English: A writing score may combine task fulfilment, organisation, language control and audience awareness. One total should be decomposed before deciding what to repair.

Mathematics: A wrong final answer may hide a correct model and minor execution error—or a fundamentally wrong representation followed by flawless calculation.

Science: Assessment may sample factual knowledge, evidence interpretation, variable control, causal explanation and transfer into unfamiliar systems. The mark distribution matters because different items are testing different scientific operations.

For parents: how should I read one bad test?

Read the paper before reading the child. Ask what the assessment sampled, where marks were lost, whether the errors cluster, what support was absent, whether timing mattered and whether the result agrees with previous evidence.

How Should Parents Read One Bad Test? provides a narrower practical route for exactly this situation.

How do we know an assessment interpretation is strong?

  • The claim matches what the assessment actually sampled.
  • The result is interpreted with its measurement limits visible.
  • Comparisons use sufficiently comparable tasks or appropriate scales.
  • Support and conditions are known.
  • Other relevant evidence is considered.
  • The conclusion remains open to correction by future performance.

The complete assessment-evidence chain

NAME THE CLAIM → CHOOSE / INSPECT THE TASK → OBSERVE PERFORMANCE → MARK OR DESCRIBE → CHECK SAMPLING AND CONDITIONS → CONSIDER RELIABILITY → INTERPRET WITHIN VALID LIMITS → COMPARE WITH OTHER EVIDENCE → DECIDE → RETEST IF THE DECISION MATTERS

Frequently asked questions

Is a school exam an accurate measure of ability?

It is evidence about performance on the content and demands sampled under those conditions. It can be very useful, but “ability” is broader than one paper and should not be inferred more strongly than the assessment supports.

Why can two tests give different results?

They may differ in difficulty, coverage, marking, timing, question format or the learner’s state. Some variation is expected. The useful task is to determine whether the difference reflects real change or changed measurement.

Are formative assessments less important because they do not count toward grades?

No. Their purpose is different. Formative evidence can be extremely valuable precisely because it arrives while teaching and learning can still adapt.

Should parents compare class averages?

Class averages provide group context but cannot tell you why an individual learner received a particular score or whether their capability improved. Use them as one signal, not the whole interpretation.

Read next

Research bridges

For professional testing principles, see the open-access Standards for Educational and Psychological Testing. For classroom evidence used to adapt teaching, see the Education Endowment Foundation overview of Embedding Formative Assessment.