Wait, what? A science news graphic summarises twelve studies of a treatment used on seedlings. Beneath the combined result it prints: “I² = 80%”. A student points at the number and says, “Easy. Eighty per cent of the studies disagree with each other.” The statement feels natural because 80% looks like a count of studies. But that is not what the statistic means.
This Reality Lab owns a narrow real-world evidence job: how to read an I² heterogeneity label on a meta-analysis or forest-plot style summary without converting it into “the percentage of studies that disagree”. The aim is not to teach Primary 5 and Primary 6 students to calculate meta-analysis statistics. The aim is to train a deeper PSLE Science habit: when a percentage appears in a scientific communication object, identify its numerator, denominator and scientific meaning before turning it into a verbal claim.
Quick answer
No. I² = 80% does not mean that 80% of the studies disagree, that 80% are wrong, or that only 20% should be trusted. I² is used in meta-analysis to describe inconsistency or heterogeneity among study results. In broad terms, it estimates the proportion of variability in the effect estimates that is associated with real between-study heterogeneity rather than ordinary sampling variation. Its interpretation depends on the size and direction of effects, the number and precision of studies, and the uncertainty around the heterogeneity estimate.
The learner habit is: do not turn a percentage statistic into a percentage of objects unless the statistic is actually counting those objects.
The exact learner job — and the non-ownership boundary
This page applies existing PSLE Science reasoning to a scientific evidence-summary label. It does not become the owner of graph reading, averages, sampling, fair testing, alternative explanations or statistical theory. It also does not teach children to conduct a medical meta-analysis. The object here is the communication move: one percentage is shown beside a combined scientific result, and a reader gives the percentage the wrong meaning.
For the underlying distinction between what a display directly reports and what a reader infers, route to How to Tell Observation, Inference, Prediction and Explanation Apart in PSLE Science. For judging whether different experimental conditions can be compared fairly, use How to Decode Variables and Fair Tests in PSLE Science Questions.
An original evidence object: twelve seedling studies
Imagine twelve independent teams testing whether a particular light schedule changes the weekly height increase of the same plant species. This is a constructed example. Each team reports an estimated effect and uncertainty. A review combines the evidence and prints:
Combined effect: positive
I² = 80%
A news caption then says: “Scientists cannot agree: 80% of studies contradict one another.”
That caption has changed the statistic into a headcount. The display did not say “9.6 out of 12 studies disagree”. It reported a heterogeneity statistic.
Observed, summarised, claimed and inferred
| Layer | What belongs there? |
|---|---|
| Observed within each study | Each team collected its own plant measurements under its own stated conditions. |
| Study result | Each study estimated an effect, with uncertainty. |
| Review summary | The meta-analysis combined results and reported I² = 80% as one description of heterogeneity. |
| Unsupported headline | “80% of the studies disagree.” |
| Reasonable next question | How different are the study effects, in what directions, under what conditions, and how uncertain is the heterogeneity estimate? |
Good evidence reading slows down at every arrow between layers.
Why a percentage can mislead even when it is calculated correctly
The symbol % often tempts us to ask, “percentage of what items?” Sometimes that is exactly right: 40% of 100 seedlings may mean 40 seedlings. But many percentages describe a relationship between quantities rather than a count of objects. Efficiency, relative change, uncertainty measures, concentration, humidity, recovery and statistical indices can all use percentages while answering different questions.
I² is one of those cases. It is not a vote. A study is not placed into a “disagrees” basket or an “agrees” basket and counted. Instead, the statistic is based on the pattern of variability among effect estimates relative to the variability expected from sampling error.
Three pictures that could all produce a worrying headline — but mean different things
Picture A: same direction, different sizes
Suppose all twelve studies estimate a positive effect, but some estimate a small increase and others a much larger increase. The studies can be heterogeneous even though they do not point in opposite directions. Saying “80% disagree” would hide that important fact.
Picture B: mixed directions
Now suppose some studies estimate a positive effect and some estimate a negative effect. This is a more serious kind of inconsistency for interpretation because even the direction varies. The same I² number cannot tell the whole story by itself; the actual effect estimates matter.
Picture C: very precise studies with small differences
When studies are highly precise, even modest differences among their estimates can contribute to a high I². That is one reason authoritative guidance warns against treating fixed numerical bands as automatic verdicts. Context matters.
The Cochrane Handbook explains that the importance of an observed I² value depends on factors including the magnitude and direction of effects and the strength of evidence for heterogeneity. See Cochrane Handbook, Chapter 10.
The representation check: inspect the forest, not only the label
A strong reader does not stare only at “I² = 80%”. Look at what the summary graphic is trying to represent:
- How many studies are included?
- Do their estimated effects point mostly in the same direction?
- Are the effect sizes similar or spread widely?
- How wide are the uncertainty intervals?
- Are the studies about the same population, conditions and outcome?
- Is the combined result sensible when the studies differ greatly?
This is evidence reading, not number worship. A single summary statistic is a signpost to inspect variation, not permission to ignore the underlying studies.
Comparison check: were the studies really asking the same question?
Imagine the twelve seedling studies differed in ways that matter:
- some used young seedlings and others mature plants;
- some used eight hours of light and others fourteen;
- some measured height after one week and others after six weeks;
- some used rich soil and others nutrient-poor soil;
- some kept temperature constant and others ran outdoors.
If results differ, the difference might reflect genuine differences in conditions rather than “bad science”. Heterogeneity can be scientifically informative because it may reveal that an effect depends on context.
Alternative explanations for variation
If study results vary, possible explanations include real biological differences, different methods, measurement error, sampling variation, different definitions of the outcome, different exposure levels, or biases in study design or reporting. One statistic cannot identify which explanation is correct.
This is a powerful transfer rule: a measure of variation is not automatically an explanation for variation.
Worked case 1: all positive, still heterogeneous
Four invented studies estimate that a treatment increases germination by 2, 5, 11 and 18 percentage points. All four estimates are positive. A high heterogeneity statistic would not mean “most studies disagree that the treatment helps”. It may mean the estimated size of the effect differs substantially across studies.
The better question becomes: why might the size differ? Seed variety? Temperature? Treatment dose? Study duration? Measurement? That leads to scientific inquiry.
Worked case 2: one dramatic outlier
Suppose eleven studies show small effects clustered near one another, while one study reports a very large effect. A headline may say “the studies are inconsistent”. But a careful review asks whether the unusual study used different conditions, had a small sample, suffered a method problem or genuinely identified a context where the effect is larger.
Rejecting the outlier automatically would also be poor science. First investigate.
Worked case 3: the same I², a different decision
Imagine two meta-analyses both report I² = 80%.
- In Review X, all effects are positive but their sizes vary.
- In Review Y, half the effects are positive and half negative.
It would be careless to treat those situations as equivalent just because the I² values match. Review Y raises a stronger question about whether one overall direction is meaningful. The statistic must be read alongside the actual pattern of results.
Tempting but invalid translations
- “I² = 80%, so 80% of studies are wrong.”
- “I² = 80%, so only 20% of the combined result is trustworthy.”
- “I² = 0%, so every study produced exactly the same result.”
- “High I² proves the treatment has no effect.”
- “Low I² proves the treatment works.”
Each statement gives I² a job it does not own.
What evidence would strengthen interpretation?
- A clear forest plot or table showing individual study effect estimates and uncertainty.
- Information about populations, methods, treatment levels and outcomes.
- An explanation of plausible sources of heterogeneity.
- Sensitivity analyses showing whether conclusions change under reasonable choices.
- A cautious conclusion that matches the consistency and direction of the evidence.
What would weaken a news-style claim?
- Reporting only I² without the study results.
- Calling I² the “percentage of studies that disagree”.
- Using one threshold as an automatic pass/fail rule.
- Ignoring opposite directions of effect.
- Combining studies with importantly different questions and then treating the average as universally applicable.
How far can the conclusion travel?
A meta-analysis summarises a defined set of studies. Its combined estimate and heterogeneity information apply to that evidence set and the question the review actually asked. They do not automatically establish what will happen in every population, every condition or every future study.
When heterogeneity is substantial, the limits of travel become especially important. A useful question is: “What conditions might change the result?”
PSLE-style transfer case
This is original practice, not an examination question.
A review combines eight experiments about a biodegradable coating and reports I² = 75%. Six experiments show the coating slows water loss, while two show little change. A website writes: “75% of scientists disagree about whether the coating works.”
Question 1: Why is the website’s statement not supported by I² alone?
Explained answer: I² describes heterogeneity among effect estimates; it is not the percentage of studies or scientists that disagree.
Question 2: What information in the scenario should be examined before deciding whether the evidence is consistent?
Explained answer: Examine the direction and size of individual effects, their uncertainty and the study conditions. Six studies showing slowing of water loss and two showing little change is different from a set split equally between opposite effects.
Question 3: Give one plausible scientific reason results might vary.
Explained answer: The experiments could differ in coating thickness, material, humidity, temperature or measurement method. Those differences should be investigated rather than labelled simply as disagreement.
Delayed independent return
- Does I² count the fraction of studies that are wrong?
- Can all studies point in the same direction and still show heterogeneity?
- Does a high I² identify the cause of variation?
- What should you inspect besides the I² number?
- What is the transferable percentage-reading habit?
Check: No; yes; no; inspect effect sizes, directions, uncertainty and study conditions; and never assume a percentage counts objects unless the quantity definition says it does.
Parent and tutor teaching guide
Do not begin by teaching a formula for I². Begin with the language trap. Put three statements on cards: “80% of plants germinated”, “efficiency = 80%”, and “I² = 80%”. Ask: “Does 80% mean the same thing in all three?” The child should discover that the symbol is not the definition.
Then draw four simple study-result arrows of different lengths. First make them all point right; then make two point left. Ask whether “difference in size” and “difference in direction” are the same scientific situation. This builds intuition for heterogeneity without turning the session into secondary-school statistics.
Authoritative sources and curriculum frame
- SEAB: 2026 PSLE Science syllabus — includes interpreting and analysing information and evaluating observations, information and methods.
- MOE: 2023 Primary Science Teaching and Learning Syllabus — supports healthy scepticism, objectivity, open-mindedness and evidence-based explanation.
- Cochrane Handbook: Analysing data and undertaking meta-analyses — explains interpretation of statistical heterogeneity and cautions that the importance of I² depends on effect magnitude and direction and the strength of evidence for heterogeneity.
The quiet habit to keep
A percentage sign is not a meaning. When science gives you “80%”, ask what relationship the number represents. If the label is I², do not turn it into a vote. Go back to the studies, inspect the pattern, ask why results differ, and let the evidence set the strength and reach of the conclusion.