PSLE-SCI-REALITY-0116
Wait, What? “We Did Not Detect a Difference” Is Not the Same Sentence as “There Is No Difference”
A product comparison tests two insulating materials. Material A keeps a model container warm for an average of 61 minutes. Material B keeps it warm for an average of 66 minutes. The study report says the difference was not statistically significant.
A headline then appears:
“Science proves the two materials work exactly the same.”
That headline says more than the result does.
A statistical test can fail to find strong enough evidence of a difference for several reasons. The true difference might really be very small. But the study might also have too few observations, measurements might vary widely, or the method might be unable to resolve a difference of the size that matters. A result that is not statistically significant therefore does not automatically prove equality.
The Reality Lab habit is: separate “absence of convincing evidence for a difference” from “convincing evidence that the two are equivalent”.
Quick Answer
- Read what the study actually tested.
- Look at the size and spread of the observed difference, not only the label “significant” or “not significant”.
- Ask whether the study had enough observations and sufficiently precise measurements to detect a difference that would matter.
- Do not translate a null result into “proven equal” unless the study used an appropriate design for demonstrating equivalence or ruling out meaningful differences.
- Keep the public conclusion narrower than the statistical result.
The Exact Learner Job This Page Owns
This page owns one real-world communication problem: a scientific report or headline turns “no statistically significant difference” into “the two conditions are the same”.
It does not turn Primary Science into a statistics course. Learners do not need to calculate p-values here. The job is evidence transfer: recognise that a test can fail to detect a difference without proving that no meaningful difference exists.
- Reality Lab Vol No.057: Does the ± Part Matter to the Claim?
- Reality Lab Vol No.033: Do the Individual Results Follow the Line?
- How to Read Repeated PSLE Science Results When Measurements Do Not Match Exactly
- How to Decide Whether an Investigation Should Measure the Whole System or a Sample
Original Reality Lab Case: Two Insulating Wraps
This is an original teaching case with constructed data. It is not copied from an examination or research paper.
Six identical model containers are wrapped with Material A and six with Material B. Each container starts at the same temperature. The time taken to cool to a fixed endpoint is recorded.
| Material A | Material B |
|---|---|
| 55 min | 57 min |
| 62 min | 63 min |
| 58 min | 70 min |
| 66 min | 69 min |
| 64 min | 72 min |
| 61 min | 65 min |
The averages differ, but the individual results also vary. A statistical test used by the fictional study does not meet its chosen evidence threshold for declaring a difference.
What can we say?
We can say the study did not obtain sufficiently strong statistical evidence, under its method and sample size, to declare a difference using that test. We cannot automatically say that A and B are physically identical, that their true average performances are exactly equal, or that no practically important difference could exist.
Observed, Claimed and Inferred
| Layer | Statement |
|---|---|
| Observed | The two sets of measurements overlap and vary. |
| Observed | The study’s chosen statistical test did not reach its significance threshold. |
| Careful statement | The study found insufficient statistical evidence of a difference under this design. |
| Overclaim | The two materials are proven to be exactly the same. |
| Missing question | Could the study have detected a difference large enough to matter? |
What “Statistically Significant” Is Trying to Do
Repeated measurements vary. If you test two groups, their averages can differ just because the particular samples happened to vary. Statistical methods help researchers judge whether the observed pattern is difficult to explain by ordinary sampling variation under a chosen model.
That is useful, but it does not create a magical border between “real” and “not real”. A significance result depends on the size of the effect, the variability of the observations, the number of observations, assumptions of the statistical model and the analysis chosen.
For Primary 5/6, the useful idea is enough: a statistical test is one piece of evidence about whether an observed difference is convincing, not a machine that proves two things equal whenever the test is negative.
The Small-Sample Problem
Imagine testing only two items from each product. Even if one product tends to perform better, two items may not provide enough information to separate the product effect from ordinary variation.
A small study can therefore produce “no statistically significant difference” simply because it has weak resolving power. This is one reason serious interpretation asks whether the study was capable of detecting a difference of the size that matters.
The High-Variability Problem
Suppose Material A usually lasts around 60 minutes and Material B around 70 minutes, but individual tests vary from 40 to 90 minutes because room conditions and sample thickness are poorly controlled. The underlying average difference may be useful, yet the noisy method can hide it.
The scientific response is not to declare equality. It is to improve the experiment, reduce uncontrolled variation, increase suitable replication or report the remaining uncertainty honestly.
The Tiny-Effect Problem Goes the Other Way
Very large studies can sometimes detect extremely small differences. A difference can be statistically convincing but too small to matter practically. That is the mirror image of today’s problem.
This tells us why strong scientific communication looks at both evidence strength and effect size. “Significant” does not automatically mean “important”, and “not significant” does not automatically mean “equal”.
What Would It Take to Support “They Are Effectively the Same”?
If the real scientific question is whether two methods are close enough to be treated as equivalent for a purpose, the study should define what size of difference would count as meaningfully different, then use a design and analysis capable of testing that question.
Nature’s statistical guidance makes this distinction explicit: an ordinary null finding should not be interpreted as evidence that no difference exists; support for equivalence requires methods designed for equivalence or other evidence appropriate to that claim.
For a Primary learner, translate that into plain language: “same enough” is its own scientific claim and deserves its own evidence.
The Representation Check: What Did the Headline Remove?
Compare these two statements:
“The study did not find statistically significant evidence of a difference.”
“Science proves there is no difference.”
The second sentence deletes the uncertainty, the study design, the sample size and the possibility that the study lacked resolving power. The transformation from report to headline changes the scientific meaning.
The Graph Check: Look at the Individual Results
A bar chart showing only two averages can hide how much the observations overlap and how variable the groups are. Whenever possible, inspect the individual measurements, uncertainty intervals or another representation that shows the spread.
This connects directly to PSLE Science habits: do not let one summary number erase the pattern in the underlying evidence.
Worked Case 1: Four Trials, Large Difference, Large Variation
Four tests of Material X give 40, 60, 70 and 90 units. Four tests of Material Y give 55, 75, 85 and 100 units. The Y values tend to be higher, but both sets vary strongly and overlap. A small study may not resolve the difference convincingly. “No significant difference” would not prove identical performance.
Worked Case 2: Many Precise Trials, Almost No Difference
Hundreds of carefully controlled tests give averages of 60.00 and 60.05 units with very little variation. A statistical analysis might detect the tiny difference. The next question becomes whether 0.05 units matters to the scientific or practical decision. Statistical detectability and practical importance are different jobs.
Worked Case 3: Two Methods Intended to Be Interchangeable
A study asks whether two thermometers can be treated as interchangeable within ±0.5°C for a school investigation. Merely failing to find a difference is not enough. The study needs evidence that likely differences are sufficiently small relative to the stated ±0.5°C equivalence goal.
Worked Case 4: “No Difference” in a Viral Graphic
A viral infographic says two plant treatments are “the same” because one small experiment did not reach statistical significance. The original data show Treatment B plants were generally taller but measurements varied widely. The fair conclusion is that this experiment did not establish a clear difference—not that biology has proven the treatments equivalent.
Tempting Reasoning That Fails
- “Not significant means no effect.” It means the chosen analysis did not produce sufficiently strong evidence for a difference under its threshold and assumptions.
- “If the averages differ, the result must be significant.” Variability and sample size also matter.
- “If the averages are close, the two are proven equal.” Close observed averages can still be uncertain.
- “A bigger sample automatically makes a result more important.” Bigger samples can detect smaller effects; importance still depends on the size and meaning of the effect.
- “Statistical significance proves causation.” Causal inference needs an appropriate experimental or observational design; a significance label alone is not a causal mechanism.
What Evidence Would Strengthen an Equality or Equivalence Claim?
- a clear definition of what size of difference would matter;
- a study large and precise enough to detect that difference;
- an analysis designed to test equivalence or otherwise bound the possible difference;
- transparent individual data or uncertainty intervals;
- repeat studies under relevant conditions;
- a result that remains close under independent measurements and alternative reasonable analyses.
What Would Weaken It?
- very few observations;
- large unexplained variation;
- wide uncertainty around the difference;
- a headline that deletes the phrase “we found insufficient evidence”;
- no definition of what difference would be scientifically meaningful;
- a study designed only to look for difference but interpreted as proof of sameness.
How Far Can the Conclusion Travel?
A well-designed study that does not detect a difference can legitimately report that result. It may narrow the range of plausible differences, especially when the study is precise and large. But the conclusion must match what the study can rule out. A weak study may leave a wide range of meaningful differences possible.
The deeper habit is calibration: the absence of strong evidence for one claim does not automatically create strong evidence for its opposite.
PSLE-Style Transfer Case
Two groups of identical model cars are tested on different tyre materials. Group A travels an average of 4.8 m. Group B travels an average of 5.2 m. The results vary widely, and the study reports “no statistically significant difference”.
A student writes: “The tyre materials definitely make no difference.”
Reasoned answer: The study did not find sufficiently strong statistical evidence of a difference, but that does not prove the tyre materials have identical effects. The sample size and variability may make a real difference difficult to detect. More precise or more extensive evidence would be needed to support a strong equality claim.
Explained Practice
Practice A: A tiny experiment finds no significant difference. What is the first caution? It may not have enough evidence to detect a meaningful difference.
Practice B: A very precise study rules out differences larger than 0.1 units, and the decision only cares about differences above 1 unit. What becomes more defensible? The two conditions may be practically equivalent for that specific purpose.
Practice C: Two articles report “significant” and “not significant” results. Can you conclude that the studies themselves significantly disagree? Not from those labels alone. Compare their effect estimates and uncertainty directly.
Delayed Independent Return: The N-U-L-L Check
- N — Number of observations: Was the study large enough for the question?
- U — Uncertainty and variability: How wide is the remaining range of plausible differences?
- L — Largest meaningful difference: What size would actually matter?
- L — Language: Does the conclusion say “insufficient evidence of difference” or overclaim “proven the same”?
Parent and Tutor Teaching Guide
Use a simple hidden-object game. Put two small objects into opaque bags and let the learner make only one noisy measurement of each bag. If the measurements look similar, ask whether the learner has proved the contents are identical. Then allow ten better measurements and discuss how stronger evidence can narrow uncertainty.
A second useful exercise is sentence repair. Give the learner: “There was no statistically significant difference, so the two treatments are the same.” Ask them to rewrite it as: “This study did not find sufficiently strong statistical evidence of a difference under these conditions.” Then ask what additional evidence would be needed to claim practical equivalence.
Authoritative Sources
- SEAB — 2026 PSLE Science Syllabus
- MOE — 2023 Primary Science Teaching and Learning Syllabus
- Nature — Statistical Guidelines, Including Interpretation of Null Results
- NIST/SEMATECH e-Handbook of Statistical Methods
Current scientific guidance is clear that absence of statistical evidence should not automatically be rewritten as evidence of absence. The official Singapore Primary Science and PSLE frame supplies the transferable learner skills: analyse information, evaluate evidence and methods, keep uncertainty visible and communicate a conclusion proportionate to what the data justify.
The Quiet Return
Science sometimes says, “We could not distinguish these clearly with the evidence we have.”
That is not the same as saying, “They are identical.”
When a headline turns “not significant” into “the same”, put the uncertainty back into the sentence.