PSLE-SCI-REALITY-0119
Wait, What? A Group Can Improve Even When Nothing Was Done to Improve It
A school science club tests twenty identical fictional light sensors under the same lamp. Most readings cluster around 100 units. A few are lower: 86, 88 and 90. The club selects only those three low readings, waits an hour, repeats the measurement and gets 94, 97 and 96.
A poster announces:
“Our adjustment improved every weak sensor.”
Perhaps the adjustment helped. But the before-and-after pattern does not prove that yet.
The three sensors were chosen because their first readings were unusually low. If ordinary measurement variation sometimes pushes a reading lower or higher than its usual value, selecting the most extreme first measurements creates a special problem: on a later measurement, some of those extremes may naturally be less extreme even if the underlying system has not changed.
This pattern is often called regression toward the mean. A Primary 5/6 learner does not need advanced statistics to use the idea. The learner needs one powerful question:
Were these cases chosen because their first results were unusually bad or unusually good?
If yes, ordinary variation can imitate improvement or decline on retest. A causal claim therefore needs a stronger comparison.
Quick Answer
- Ask how the group was selected. Were only the lowest, highest, worst or best first results chosen?
- Keep the individual first results visible. Selection based on an extreme first measurement matters.
- Ask whether the measured quantity naturally varies from one reading to another.
- Find out whether an intervention happened between measurements and whether a suitable comparison group was measured over the same period.
- Retest comparable units that were not selected only because of an extreme first result, or obtain repeated baseline measurements before the intervention.
- Do not call ordinary movement away from an extreme result proof that the treatment caused improvement.
The Exact Learner Job This Page Owns
This page owns one real-world evidence-transfer job: evaluating a before-and-after claim when follow-up measurements were performed only on cases selected because their first measurements were unusually high or low.
It does not become a general statistics lesson. It does not replace the PSLE Science owners of repeated trials, fair comparisons, variables, causation or anomalous results. It applies those familiar skills to a subtle communication pattern that can make an intervention look effective even when selection and ordinary variation explain part of the apparent change.
- Reality Lab Vol No.004: “X Causes Y” — What Evidence Would a Causal Headline Need?
- Reality Lab Vol No.022: Are the Before-and-After Pictures Really Comparable?
- Reality Lab Vol No.029: What Happened to the Missing Results?
- Reality Lab Vol No.052: What Was Actually Random?
Original Reality Lab Case: The “Weakest Five” Repair Test
This is an original composite teaching case with constructed data. It is not copied from an examination, paper or commercial study.
Twenty fictional devices are tested twice under the same stated conditions. Their true performance is intended to be stable over the short interval, but each measurement contains a small amount of ordinary variation from positioning, temperature and instrument noise.
After the first test, a team chooses the five lowest results for a “repair check”. The devices receive a harmless sticker that has no physical effect. One hour later the team measures only those five again.
| Selected device | First reading | Second reading | Change |
|---|---|---|---|
| A | 86 | 95 | +9 |
| B | 88 | 94 | +6 |
| C | 89 | 97 | +8 |
| D | 90 | 93 | +3 |
| E | 91 | 96 | +5 |
The result looks dramatic: all five selected devices improved.
But the sticker cannot explain the change because it has no physical mechanism. What happened? The first test selected devices whose readings happened to lie at the low end of the observed spread. On a new measurement, ordinary variation did not have to push the same devices downward by the same amount again. Their second readings therefore tended to be less extreme.
This does not mean every extreme result must move inward on retest. Some can become even more extreme. It means that when selection is based on an unusually extreme first measurement, movement toward more typical values can arise without the claimed intervention being the cause.
Observed, Claimed and Inferred
| Layer | Statement |
|---|---|
| Observed | The five selected devices had low first readings and higher second readings. |
| Selection fact | The devices were selected specifically because their first readings were among the lowest. |
| Public claim | “The intervention caused the improvement.” |
| Alternative explanation | Ordinary measurement or system variation made extreme first values less extreme on retest. |
| Evidence needed | A comparison design that can separate intervention effect from selection plus ordinary variation. |
Why Selecting the Extremes Changes the Evidence Problem
Suppose a measurement is usually close to 100 but can move a few units up or down from one reading to another. A reading of 87 is therefore unusual. If you choose the device because it happened to show 87, you have deliberately selected a case containing an extreme measurement.
On the next test, there is no reason the same random influences must line up in exactly the same direction. The second result may be closer to the device’s usual level.
The selection rule and the natural variability are therefore connected. This is what makes the problem different from an ordinary before-and-after comparison of a group chosen independently of its first result.
Do Not Confuse This With “The First Measurement Was Wrong”
An extreme first result can be a valid observation. Regression toward the mean does not require the first result to be a mistake. It only requires some variation between repeated measurements and selection based on an extreme first value.
The correct scientific response is not to erase the first result. It is to recognise that a single extreme measurement may contain both a real underlying difference and temporary variation.
The Selection Check: Who Was Allowed Into the Before-and-After Story?
A convincing graphic may show ten “weak performers” before a programme and the same ten after it. The missing question is how those ten were chosen.
- Were they chosen because their first score was below a threshold?
- Were they the lowest ten out of a much larger group?
- Were all tested units followed up, or only the extremes?
- Was the threshold chosen before seeing the data?
- Would a unit with an ordinary first result have been included?
Selection rules are part of the method. If the public claim hides them, the before-and-after picture may look more causal than the evidence deserves.
The Baseline Check: Was One Measurement Enough?
One way to reduce the problem is to understand the baseline better before acting. If a device is measured several times before an intervention, we can see whether one low reading is typical or unusual for that device.
Imagine a device gives 98, 99, 88, 100 and 97 on five repeated baseline tests. If it is selected for repair only because of the 88, the surrounding measurements reveal that 88 may not describe its usual state.
Repeated baseline evidence does not solve every causal question, but it prevents one unusually extreme reading from quietly becoming the entire diagnosis.
The Comparison Check: What Happened to Similar Extreme Cases Without the Intervention?
A particularly useful comparison is another group selected by the same rule but not exposed to the intervention. If both groups move toward more typical values, selection plus ordinary variation becomes a strong alternative explanation.
If the treated group improves much more than a comparable untreated group under otherwise similar conditions, the evidence for a treatment effect becomes stronger.
The exact design depends on the scientific question. The Primary-level habit is simply to ask for the comparison that can separate the claimed cause from the obvious alternative.
The Representation Check: A Before-and-After Arrow Can Hide the Selection Rule
An infographic may show a red bar labelled “before: 70” and a green bar labelled “after: 85”. A large upward arrow makes improvement visually obvious.
But the image may not show that the group was created by selecting everyone below 75 on the first test. The group therefore starts unusually low by design.
A more honest graphic would state the selection rule, show the distribution of first measurements, preserve individual follow-up results and include a suitable comparison if a causal claim is being made.
Regression Toward the Mean Does Not Mean “Everything Becomes Average”
The phrase can be misleading if taken literally. It does not say that every unusual result must become ordinary, that true differences disappear, or that all systems move toward one common value.
It describes a statistical tendency that appears when measurements contain variability and cases are selected because an initial measurement is extreme. The next measurement often contains a less extreme combination of temporary influences.
A real intervention can also work at the same time. The scientific problem is to separate how much change comes from the intervention from how much could occur without it.
Alternative Explanation 1: Ordinary Measurement Variation
Small changes in instrument reading, object position, environmental conditions or observer judgement can make repeated measurements differ. If the first value is extreme partly because several small influences pushed in the same direction, the second value may be less extreme when those influences change.
Alternative Explanation 2: The System Naturally Fluctuates
Some real systems vary through time even when measurement is perfect. Temperature, plant growth rate, brightness, water flow and many biological or environmental quantities can rise and fall. Selecting a moment of unusually low performance and measuring later can therefore create apparent recovery without an intervention causing it.
Alternative Explanation 3: Conditions Changed for Everyone
Perhaps the room warmed, the lamp stabilised, the operator became more practised or the power supply changed. If only the selected low group is retested, these common changes may be mistaken for an intervention effect.
Alternative Explanation 4: The Intervention Really Helped
This possibility remains alive. Regression toward the mean is not a magic argument against every improvement claim. It is a reason to design the comparison carefully. Strong evidence can still show that an intervention adds improvement beyond what selection and ordinary variation would predict.
What Evidence Would Strengthen the Causal Claim?
- The selection rule is stated clearly.
- Repeated baseline measurements show whether the first extreme result is stable.
- A comparable group selected using the same threshold is followed without the intervention where appropriate.
- All selected cases are reported, including those that do not improve.
- The intervention is defined before outcomes are known.
- Measurement conditions are kept comparable.
- The effect remains on later independent measurements rather than disappearing immediately.
- The result is reproduced in a new group rather than relying only on the first extreme-selected sample.
What Would Weaken It?
- Only the worst first results are retested.
- The selection rule is hidden.
- There is no comparison group or repeated baseline.
- Cases that fail to improve disappear from the report.
- The second measurement uses easier conditions.
- The intervention is changed after the first result is seen.
- The headline says “caused” when the evidence only shows “changed after”.
Worked Case 1: The Lowest Five Thermometers
Thirty thermometers measure the same stable water bath. The five lowest readings are retested. Four rise on the second measurement. Does this prove the thermometers repaired themselves? No. They were selected precisely because the first readings were unusually low. Measurement variation alone can make the retest look better.
Worked Case 2: The Shortest Seedlings
A class chooses the five shortest seedlings on Monday, gives them a harmless coloured label and measures only those five again on Friday. They have grown. Can the label be credited? No. Seedlings grow over time, and the selected group also began at an extreme. A comparison with similarly selected unlabelled seedlings and proper control of conditions would be needed to test the label claim.
Worked Case 3: The Highest Noise Readings
A monitoring team chooses the four locations with the highest one-minute noise readings and returns the next day. All four are quieter. That does not prove a new sign reduced noise unless the sign was actually introduced under a design that distinguishes its effect from ordinary time-to-time variation and the extreme-selection problem.
Worked Case 4: Improvement That Survives the Stronger Test
Two sets of low-performing fictional devices are selected using the same rule. One set receives a real adjustment; the other does not. Both improve slightly on retest, but the adjusted group improves substantially more, and the difference remains on several later measurements under changed conditions. That design gives stronger evidence that the adjustment contributed beyond ordinary movement away from an extreme baseline.
Tempting Reasoning That Fails
- “They improved after the intervention, so the intervention caused it.” Time order alone does not rule out selection and ordinary variation.
- “Regression to the mean proves the treatment did nothing.” No. It is an alternative explanation that good design must separate from a real effect.
- “The first extreme result was false.” Not necessarily. It may be a genuine but variable measurement.
- “If everyone improves, causation is proven.” If everyone was selected for an extreme first value, common movement can still occur without the intervention being the cause.
- “Just average the before and after values.” Averaging does not repair a biased selection design.
Model and Measurement Limits
How much regression toward the mean occurs depends on how variable the measurement or system is and how extreme the selection rule is. A very stable measurement with little variation may show little effect. A highly variable quantity selected at an extreme can show a larger apparent return toward typical values.
Primary learners do not need to calculate this effect. They need to recognise when the study design makes it plausible.
How Far Can the Conclusion Travel?
If an extreme-selected group improves on retest, the safe first conclusion is: the selected group had less extreme measurements later. To say the intervention caused the change requires evidence that distinguishes the intervention from ordinary variability, time effects, changed conditions and selection.
This disciplined wording is not weakness. It keeps the claim matched to the evidence.
PSLE-Style Transfer Case
A pupil measures the bounce height of twenty supposedly similar balls. The three lowest first measurements are 42 cm, 43 cm and 44 cm. The pupil rubs those three balls with a dry cloth and measures them again. The new heights are 47 cm, 46 cm and 48 cm. The pupil concludes that rubbing increased bounce height.
Question: Give one reason the before-and-after evidence alone does not prove that rubbing caused the increase.
Reasoned answer: The balls were selected because their first measurements were unusually low. Ordinary variation in repeated bounce measurements could make their later results less extreme even without rubbing causing a change. A suitable comparison or repeated baseline is needed.
Explained Practice
Practice A: A study retests every unit, not only the extremes. Does regression toward the mean disappear completely? The special selection problem is reduced because follow-up was not restricted to extreme first values, although repeated measurements can still vary.
Practice B: A group is selected because its first values are extremely high, then later values are lower. Can the same reasoning apply? Yes. Extremes can move toward more typical values from either direction.
Practice C: The treated extreme group improves more than a comparable extreme group that received no treatment, and the difference persists on later tests. What changes? The intervention explanation becomes stronger because an important alternative has been tested rather than ignored.
Delayed Independent Return: Three Questions to Ask Tomorrow
- Why were these cases chosen? If the answer is “because the first result was extreme”, selection matters.
- How variable is the measurement or system? If repeated values naturally move, one extreme baseline deserves caution.
- What comparison separates ordinary return from a real intervention effect? Look for repeated baselines, comparable controls and later independent checks.
Parent and Tutor Teaching Guide
Use a simple repeated-measurement game rather than beginning with the term “regression to the mean”. Ask a learner to throw a soft ball toward a target ten times. Choose only the two worst throws. Then repeat just those two positions or attempts under similar conditions. If the next attempts are better, ask whether a new teaching method caused the change—or whether choosing the worst first attempts made improvement more likely on repetition.
Then strengthen the design. Repeat all attempts, or create a comparison group selected by the same rule. The learner should see that the key issue is not the vocabulary. It is whether the study can separate the claimed cause from the way the group was chosen.
For a more scientific version, use repeated readings of the same stable object or simulated sensor values with small random variation. The goal is to make selection visible without teaching advanced formulas.
Authoritative Sources
- Singapore Examinations and Assessment Board — 2026 PSLE Science Syllabus
- Ministry of Education, Singapore — 2023 Primary Science Teaching and Learning Syllabus
- Barnett, van der Pols & Dobson — Regression to the mean: what it is and how to deal with it
The review above explains that regression to the mean can occur when repeated measurements vary and study participants or cases are selected because their baseline values are extreme. For PSLE Science, the transferable habit is fully compatible with the official inquiry frame: evaluate the method, keep alternative explanations alive and avoid making a causal conclusion from a before-and-after pattern that the design cannot uniquely explain.
The Quiet Return
The most impressive improvement can begin with a hidden choice.
If you choose a group because its first measurement is unusually bad, the second measurement has inherited that selection problem before the intervention even begins.
When extreme results improve, ask whether the cause changed—or whether the measurement simply stopped being so extreme.