Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

PSLE Science Reality Lab Vol No.053 | “One Test Worked” — How Many Other Tests Were Tried?

PSLE-SCI-REALITY-0053

Wait, What? One amazing result can become less amazing when you discover there were 40 chances to find one.

A fictional science post announces:

“New coating reduces water loss by 34%!”

The number is real inside the fictional dataset. One comparison really did show a 34% difference. Then you read the method and discover that the researchers tested five coatings, at four temperatures, using two drying times. That creates many possible comparisons. Most differences were small. One combination produced the dramatic 34% result, and that combination became the headline.

Has the 34% result become fake? No. Has it become weaker evidence than the headline first suggested? Possibly. The answer depends on what was planned before the data were seen, how many comparisons were made, whether the standout result repeats, and whether a scientific mechanism explains why that particular condition should behave differently.

This is an advanced scientific habit, but the core idea is Primary-friendly: the more places you look for something unusual, the less surprising it is that one unusual-looking result may appear by chance.

Quick Answer

When one impressive result is highlighted, ask how many outcomes, conditions, time points, groups or comparisons were examined. Then ask whether the headline target was chosen before the results were known or selected after researchers saw which comparison looked strongest. A result selected from many opportunities can still be real and important, but it needs stronger follow-up: repetition, independent confirmation, a pre-declared test, and evidence that the effect survives new data.

Reality Lab rule: Do not inspect only the winning result. Inspect the size of the search that produced the winner.

Owned Learner Job

This article owns one transfer job: how to evaluate a scientific claim built around one standout result when many possible tests or comparisons were tried. It is not a general statistics lesson and it does not teach formal significance testing. It also differs from Reality Lab Vol No.009, which asks whether failed trials were omitted. Here, the full set of trials may exist; the problem is that one unusually favourable comparison is selected as though it were the only question that was asked.

It also connects to Reality Lab Vol No.035, which asks whether a prediction was written before the results were known. Pre-declaring the scientific question can help distinguish a genuine test of a prior idea from a pattern discovered only after searching through many possibilities.

The Original Reality Lab Case: Twenty Ways to Win

A fictional material company compares Coating Q with an untreated control. It measures water loss after 30 minutes under five humidity conditions and at four temperatures. That creates 20 condition combinations.

The original teaching data look like this:

Condition groupTypical observed difference between Q and control
Most of the 20 conditions0% to 8% lower water loss with Q
Three conditions9% to 15% lower water loss with Q
One condition34% lower water loss with Q

These are original illustrative data, not results from a real commercial product.

The advertisement prints only the largest difference: “34% lower water loss.”

A learner should not instantly reject the number. Instead, reconstruct its position inside the experiment. Was 34% the answer to a question specified before testing? Or did researchers search the 20 conditions and then choose the largest difference? Does the effect repeat when the same condition is tested again with new specimens? Does a plausible mechanism explain why only that condition shows such a large change?

Those questions turn a headline into an evidence investigation.

Why More Searches Create More Chances for a Standout

Imagine rolling one ordinary die once. Getting a six is noticeable. Now roll the die 30 times. Seeing at least one six would no longer surprise you. Nothing dishonest happened. The number of opportunities changed what counts as surprising.

Scientific datasets can create a similar problem. If you compare many groups, many outcomes, many times or many subgroups, ordinary variation may produce one result that looks unusually large. The more comparisons you inspect, the more opportunities you have to discover an extreme-looking result.

This does not mean that every unusual result found in a large dataset is false. It means the result deserves a different evidential status: often interesting finding that needs confirmation rather than final proof.

One Question Before the Results vs One Winner After the Results

Compare two investigations.

Investigation A: Target chosen first

Before collecting data, the researchers write: “We predict Coating Q will reduce water loss at 35°C and 60% relative humidity because its proposed mechanism should become important under those conditions.” They define the comparison and method in advance, perform the experiment, and observe a large difference.

The result directly tests a prior prediction.

Investigation B: Winner chosen later

The researchers measure 20 conditions without a single declared primary prediction. Afterward, they inspect all results, find the 34% difference, and build the headline around it.

The 34% observation is still evidence. But it has also helped choose the scientific question after the fact. A clean next step is to treat that condition as a new hypothesis and test it again independently.

This is why preregistration and other prospective planning practices can be useful in research: they help preserve a record of what researchers intended to test before seeing the outcome. They do not make a study automatically correct, but they can make the evidence path more transparent.

How Many “Chances to Win” Were There?

You do not need advanced statistics to count scientific opportunities. Ask whether the researchers examined:

  • many different products or treatments;
  • many temperatures or environmental conditions;
  • many measurement times;
  • many outcomes from the same experiment;
  • many subgroups;
  • many graph cut-offs or time windows;
  • many different calculations;
  • many possible pairwise comparisons.

Each extra choice can create another place for an unusually favourable-looking result to appear. The exact mathematical treatment of multiple comparisons belongs to higher-level statistics. The Primary Science habit is simpler: count the search before admiring the winner.

This Is Not the Same as Hiding Failed Trials

Reality Lab Vol No.009 asks whether failed trials were left out. That is a missing-evidence problem. Vol No.053 can happen even when every trial remains available.

Suppose a report honestly publishes all 20 comparisons in an appendix but its press release says only, “34% improvement.” The evidence has not disappeared. The communication has selected the most dramatic part of a larger result set.

A careful reader should restore the context before interpreting the magnitude.

This Is Not the Same as an Anomalous Result

An anomalous result is an observation that differs markedly from nearby or repeated results and may deserve investigation. A selected winner may be an anomaly—or it may be a genuine effect that occurs only under one condition. You cannot decide merely from its size.

The next scientific move is discriminatory: repeat the condition, inspect the method, test the proposed mechanism, and compare it with neighbouring conditions.

Worked Case 1: Ten Fertilisers, One Spectacular Plant

A fictional school garden compares ten nutrient mixtures, each on several similar plants. Most groups are close. Mixture H has one plant that grows exceptionally tall, and a poster says, “Mixture H creates giant growth.”

Better reasoning: Ask whether the headline is based on one plant, the group average, repeated plants, and whether H performs similarly in a new experiment. The more plants and mixtures examined, the more chances there are to encounter one extreme individual. One spectacular plant can generate a useful hypothesis without yet establishing a general rule.

Worked Case 2: Twelve Time Points, One Perfect Moment

A product is measured every hour for twelve hours. At hour 7, the difference from the control is largest. The advertisement reports only hour 7.

This is not automatically wrong. Perhaps hour 7 was chosen in advance because the mechanism predicts a peak there. But if hour 7 became important only after researchers scanned all twelve time points, the standout needs confirmation. This connects to Reality Lab Vol No.007 on selective time windows.

Worked Case 3: Twenty Outcomes From One Device Test

A fictional device is tested for speed, temperature, noise, energy use, vibration, wear, pressure, brightness and several other outcomes. One outcome improves dramatically, two worsen, and most barely change. The headline says, “Scientific testing shows major improvement.”

The learner should ask which outcome mattered before testing, which outcomes are relevant to the product’s claimed job, and whether the headline represents the full evidence. One selected improvement may not summarise the device as a whole.

Worked Case 4: The Result Repeats

Now suppose the 34% condition is tested again on a new batch using the same pre-specified method. The second experiment shows a 31% difference. A third independent group later reports a similar effect under comparable conditions.

The evidence has changed substantially. The original result began as a standout among many possibilities. Repetition with new evidence makes the pattern harder to explain as a one-off winner. This is how an interesting finding can mature into stronger evidence.

The Selected-Winner Audit

  1. Name the headline result. What exact comparison produced it?
  2. Count the search space. How many conditions, outcomes, groups or time points were examined?
  3. Check timing of the question. Was this comparison specified before the results were known?
  4. Inspect neighbouring results. Is the winner part of a consistent pattern or isolated?
  5. Check the mechanism. Is there a scientific reason this condition should differ?
  6. Demand a new test. Does the result repeat with fresh specimens or data?
  7. Limit the conclusion. Until confirmation, describe it as a finding under a particular condition rather than a universal rule.

What Evidence Would Strengthen the Standout Result?

  • the key comparison defined before data collection;
  • a mechanism that predicts why the effect should occur there;
  • transparent reporting of all relevant tested outcomes and conditions;
  • the standout surviving repeated trials;
  • the effect appearing in a new batch or sample;
  • independent researchers finding a similar result;
  • neighbouring conditions forming a scientifically coherent pattern rather than one isolated spike.

What Would Weaken It?

  • the headline comparison chosen only after searching many possibilities;
  • many outcomes tested but only one favourable result discussed;
  • the standout disappearing when repeated;
  • no mechanism explaining why that condition should differ;
  • neighbouring results showing no related pattern;
  • changing the calculation or cut-off repeatedly until a dramatic result appears;
  • presenting an exploratory finding as though it were a pre-planned prediction.

Tempting Reasoning That Fails

  • “The biggest result must be the truest result.” Magnitude alone does not establish reliability.
  • “If many tests were done, every result is meaningless.” Too strong. Multiple testing changes how cautiously a standout should be interpreted; it does not erase all evidence.
  • “A pattern found after looking at the data is worthless.” No. It can generate a valuable new hypothesis. It simply needs a fresh test.
  • “The result repeated in the same dataset, so it is confirmed.” Re-analysing the same data is not the same as obtaining new independent evidence.
  • “The headline gives a precise percentage, so the target must have been planned.” Precision of wording does not reveal when the question was chosen.

PSLE-Style Transfer Case: Five Materials, Four Conditions

A learner compares five insulating materials at four different starting temperatures. Most comparisons show small differences. Material E at the highest starting temperature shows a much larger difference than all other cases.

Which next investigation best tests whether the standout is a reliable scientific effect?

  1. Repeat only the easiest low-temperature condition.
  2. Repeat Material E at the high-temperature condition using new comparable specimens and the same valid method.
  3. Delete the other 19 comparisons.
  4. Assume the largest difference is correct because it is the largest.

Answer: 2. The standout has generated a specific new question. The cleanest confirmation is a fresh test targeted at that condition.

Practice 1: One Winner in 30

A report tests 30 combinations and highlights the only one with a very large improvement. What should you ask before calling it a general effect?

Explained answer: Ask whether that comparison was chosen in advance, whether all 30 results are visible, whether the large result repeats, and whether a scientific mechanism predicts it. The winner should be interpreted in the context of the 29 other opportunities.

Practice 2: Pre-Planned Target

A team writes its target comparison before collecting data and later finds a large result exactly there. Does that make the conclusion automatically correct?

Explained answer: No. Planning strengthens the interpretation by separating prediction from after-the-fact selection, but measurement quality, controls, sample size, alternative explanations and replication still matter.

Practice 3: Same Dataset, New Analysis

A researcher finds a surprising pattern, then uses the same dataset three more ways and keeps finding related versions of the pattern. Is that the same as three independent replications?

Explained answer: No. The analyses may add understanding, but they share the same underlying evidence. Independent replication requires new evidence that is not merely another view of the same dataset.

Delayed Independent Return

The next time a headline presents one striking percentage, train yourself to ask three quiet questions before reacting:

  1. How many possible results could have become the headline?
  2. Was this target chosen before or after the data were seen?
  3. Has the standout survived a genuinely new test?

Those three questions do not make you cynical. They make you harder to impress for the wrong reason.

Routes to Existing PSLE Science Skills

Parent and Tutor Teaching Guide

Use a simple classroom game. Hide one star card among 20 ordinary cards. Let the learner draw one card. Finding the star on the first draw feels remarkable. Now let them draw all 20 cards. Finding the star somewhere is guaranteed. The point is not probability calculation. It is to show that the number of opportunities changes how we interpret a standout.

Then move to an original scientific dataset with several conditions. Ask the learner to circle the largest difference. Next ask, “Did we decide to care about this condition before we saw the table, or because it became the biggest number?” That single question exposes the distinction between prediction and exploration.

Do not teach the child to dismiss exploratory science. Exploration is how many important questions are discovered. The repair is to change the language: We found something interesting; now let us test whether it comes back.

The learner is ready when they can admire an unusual result and still ask for its search context and independent return.

Authoritative Sources

The Quiet Return

Science often begins by noticing something strange. The discipline comes next.

When one result shines brighter than the rest, do not extinguish it and do not worship it. Ask how many places researchers looked, whether the target existed before the result, and whether reality produces the same answer when the question is asked again.

A winner becomes knowledge not because it won once, but because it survives another fair encounter with the world.