Wait, What? A Bigger Difference Is Not Automatically Better Evidence
Two PSLE Science investigations can produce very different-looking results. One shows a dramatic difference between set-ups but uses a weak comparison and only one trial. Another shows a smaller difference but uses a fairer method, suitable measurements and repeated evidence.
Which gives stronger evidence?
Not necessarily the one with the bigger numerical difference. Effect magnitude and evidence strength are different scientific jobs.
HOW BIG IS THE EFFECT? ≠ HOW STRONGLY DOES THE EVIDENCE SUPPORT THE CLAIM?
Quick Answer
Separate the questions:
FIRST ASK WHAT WAS OBSERVED AND HOW LARGE THE DIFFERENCE OR CHANGE IS → THEN ASK HOW TRUSTWORTHY THE COMPARISON IS → CHECK FAIRNESS, MEASUREMENT, REPEATS, VARIATION, RANGE AND CLAIM ALIGNMENT → DO NOT USE EFFECT SIZE AS A SUBSTITUTE FOR EVIDENCE QUALITY → DO NOT USE REPEATABILITY AS A SUBSTITUTE FOR EFFECT SIZE.
The Exact PSLE Science Learning Job This Guide Owns
This guide owns one learner job: distinguishing the magnitude of an observed scientific effect from the strength of the evidence supporting a claim, so learners do not confuse “larger result” with “better evidence”.
The existing guide How to Decide Which PSLE Science Investigation Gives Stronger Evidence for a Claim owns the broader comparison of evidence quality across investigations. This page owns the specific confusion between evidence strength and effect magnitude.
Why This Matters for Current PSLE Science
For examination from 2026, PSLE Science assesses knowledge with understanding and application of knowledge and scientific inquiry, including interpreting and analysing information, evaluating observations, information and methods, and communicating explanations and reasoning. These jobs require learners to judge both what the result says and how securely the method supports the conclusion.
The current Primary Science syllabus develops learners through connected themes and inquiry practices. This guide stays within that learning role: it does not introduce formal statistical effect sizes or advanced significance testing. It teaches a Primary-level distinction between size of observed change and quality of evidence.
Two Questions, Two Answers
| Question | What to inspect | Example answer type |
|---|---|---|
| How big is the effect? | Difference, change, rate, amount, direction, range | “Set-Up A increased by more than Set-Up B.” |
| How strong is the evidence? | Fair comparison, measurement relevance, repeats, consistency, control, specimen variation, method limitations | “Investigation 2 provides stronger support because the comparison is fairer and the result was reproduced.” |
Worked Example 1 — Large Difference, Weak Evidence
Imagine Set-Up A gives a result of 80 units and Set-Up B gives 20 units. That is a large observed difference. But the specimens were different sizes at the start, the temperature also differed and only one trial was performed.
The learner may correctly say the observed difference is large. The learner should not say the evidence strongly proves the tested condition caused that difference. Several other relevant differences remain.
Worked Example 2 — Small Difference, Stronger Evidence
Another investigation produces 52 units versus 48 units. The difference is smaller. However, the specimens are comparable, the changed condition is isolated, the measurement is appropriate and repeated trials give similar results.
This may be stronger evidence for a genuine relationship than the dramatic 80-versus-20 comparison, even though the observed effect is smaller.
But there is one more check: can the measuring method actually distinguish 52 from 48? If the instrument’s resolution is too coarse, the apparent small difference may not be supported. Evidence strength depends on the whole measurement system.
Worked Example 3 — Repeating a Result Does Not Make the Effect Larger
A learner repeats a fair investigation five times and gets similar differences each time. The repeated evidence increases confidence that the observed pattern is consistent under those conditions.
It does not make the effect itself five times larger. Repetition strengthens the evidence base; it does not multiply the scientific outcome.
Worked Example 4 — A Bigger Effect Can Be Easier to Detect
Large effects can sometimes be easier to distinguish from measurement variation. That can make the evidence easier to interpret. But “easy to see” is still not the same as “well established”. If the design is unfair or the wrong quantity is measured, a dramatic difference may still support the wrong explanation.
Worked Example 5 — Strong Evidence Can Support “No Large Effect”
Suppose a well-designed investigation repeatedly finds almost no measurable difference across the tested conditions, using a method sensitive enough to detect the kind of difference expected.
Strong evidence does not always support a big effect. It can instead support the narrower conclusion that no large effect was detected under these conditions. The exact wording should remain within the method’s resolution and tested range.
What Makes Evidence Stronger?
- the investigation actually tests the stated claim;
- relevant conditions are controlled or matched;
- the measured outcome answers the question;
- the instrument has suitable range and resolution;
- results are repeated or supported by enough comparable evidence;
- specimen variation is handled sensibly;
- the conclusion stays within what was tested;
- alternative explanations are reduced by the design.
None of those points says the numerical effect must be large.
What Describes Effect Magnitude?
- how much a measured quantity changed;
- how far apart comparable results are;
- how much faster or slower one case is;
- how large a difference remains after a fair comparison;
- whether the observed change is large enough for the measurement method to distinguish.
Effect magnitude should always be attached to the actual scientific quantity and comparison. “A big effect” without saying what changed is incomplete.
The PSLE Science Reasoning Chain
READ GIVEN INFORMATION → IDENTIFY THE SCIENTIFIC OBJECT OR RELATIONSHIP → DESCRIBE THE OBSERVED DIFFERENCE → CHECK THE QUANTITY AND SCALE → SEPARATE EFFECT MAGNITUDE FROM EVIDENCE QUALITY → EVALUATE METHOD AND REPEATS → SELECT THE RELEVANT CONCEPT → EXPLAIN THE MECHANISM WHERE REQUIRED → STATE THE CLAIM AT THE STRENGTH THE EVIDENCE SUPPORTS.
Observable Failure Signatures
| Failure signature | Likely weak link |
|---|---|
| “The difference is huge, so the experiment proves it.” | Effect magnitude substituted for evidence quality |
| “Five repeats means the effect is five times stronger.” | Evidence amount confused with effect size |
| A tiny numerical difference is accepted although the scale cannot distinguish it | Measurement resolution ignored |
| A fair repeated investigation is dismissed because the difference looks small | Visual drama mistaken for evidence strength |
| A large difference from unmatched starting conditions is explained causally | Comparison fairness ignored |
| “Stronger evidence” is used without naming why the evidence is stronger | Evaluation criteria missing |
Find the Earliest Weak Link
- What exact quantity was measured?
- How large is the observed difference or change?
- Can the measurement method resolve that difference?
- Were the compared cases scientifically fair?
- Were relevant starting conditions comparable?
- Was the result repeated or independently supported?
- Could another condition explain the difference?
- What claim does the evidence actually support?
- Am I using “large” to describe the result or the evidence?
Misconception Repair — “Bigger Difference Means Stronger Proof”
No. A large observed difference can be produced by the tested condition, a confound, a starting difference, a measurement problem or several factors at once. Evidence becomes stronger when the method makes those alternatives less plausible and the measurement is trustworthy.
Misconception Repair — “More Repeats Mean a Bigger Effect”
Repeats help show whether a result is consistent. They strengthen the evidence for the observed effect. They do not change the amount of the effect already measured in each comparable trial.
Misconception Repair — “Small Effects Do Not Matter”
The scientific question determines what matters. A small but consistently measured difference can be important evidence. A dramatic difference can be scientifically uninformative if the comparison is invalid.
Practice Protocol: Size Column, Strength Column
For each investigation you practise, make two short columns:
| Effect magnitude | Evidence strength |
|---|---|
| What changed? | Was the comparison fair? |
| By how much? | Was the right quantity measured? |
| In which direction? | Could the instrument resolve it? |
| Across which conditions? | Were results repeated or supported? |
| Within what tested range? | What alternative explanations remain? |
Then write two separate sentences: one about the size of the observed result, one about the quality of the evidence. Only after that decide what scientific explanation or conclusion is justified.
Unfamiliar Transfer Challenge
Create two original investigations about the same relationship:
- Investigation A has a large difference but one obvious confound and one trial.
- Investigation B has a smaller difference but a fair comparison, suitable resolution and repeated results.
Answer four questions:
- Which has the larger observed effect?
- Which gives stronger evidence for the claim?
- Why are those not the same question?
- What extra evidence would strengthen the weaker investigation?
Delayed Independent Return
Three to five days later, use a fresh pair of investigation summaries. Before reading any model answer, write SIZE beside statements about how much changed and STRENGTH beside statements about method quality, repeatability or support. Then explain which claim is best supported.
Evidence-vs-Effect Receipt
- I know what quantity the effect refers to.
- I can describe how large the observed difference is.
- I check whether the measurement can resolve the difference.
- I evaluate fairness separately from effect size.
- I know what repeats strengthen.
- I do not multiply an effect because there are more trials.
- I do not call a dramatic result strong evidence without checking the method.
- I match the final claim to both the observed effect and the quality of the evidence.
Parent and Tutor Teaching Guide
Give the learner two deliberately contrasting investigation summaries: one dramatic but weakly controlled, one modest but well designed. Ask two separate questions: “Which result is bigger?” and “Which evidence is stronger?” Do not allow one answer to substitute for the other.
Then ask the learner to name the reason for stronger evidence: fair comparison, suitable measurement, repeated consistency, better control, more representative sampling or another feature actually present in the scenario. Avoid vague praise such as “more reliable” unless the learner can explain what makes it so.
Finally, give a small difference near the instrument’s resolution. This forces the learner to connect magnitude and measurement without collapsing them into the same concept.
Useful Internal Routes
- PSLE Science Learning Guide
- Decide which investigation gives stronger evidence
- Decide whether a small difference is distinguishable at the scale
- Tell repeatability from fair-test quality
- Return to the newest PSLE Science guides
Authoritative References
- SEAB — PSLE Science syllabus, for examination from 2026
- MOE — Science Teaching & Learning Syllabus, Primary, 2023
Evidence and Boundary Note
This guide uses “effect magnitude” in a simple Primary-level sense: how large an observed change or difference is. It does not teach formal statistical effect-size calculations or significance testing, and it does not claim such techniques are PSLE requirements. The learner job is to keep result size separate from evidence quality.
The Quiet Return
Science is not impressed by drama.
A big result can come from a bad comparison. A small result can survive a very careful test. First ask what changed. Then ask why you should trust the evidence. Keep those two questions separate, and your conclusions become much harder to fool.