PSLE-SCI-REALITY-0068
Wait, What? Two careful scientists can use the same data, make different reasonable analysis choices and get results that are not identical.
That sounds uncomfortable. If science is objective, should there not be one correct button to press and one unavoidable answer?
Sometimes there is. If you are adding four measured masses, arithmetic fixes the result. But many real investigations contain justified choices: which time window to examine, how to summarise variation, how to define an event, how to deal with a clearly invalid measurement, or which reasonable model to use for a complicated pattern.
A scientific claim described as robust is often claiming something stronger than “we got this answer once”. It is claiming that the main conclusion survives other reasonable ways of analysing the same evidence.
That does not mean every possible analysis must agree. It does not mean the study is automatically true. And it is not the same as another team collecting new data.
This Reality Lab teaches the missing question: if we change a reasonable analysis choice without changing the scientific question, does the conclusion still stand?
Quick Answer
A robustness check asks whether the same underlying data lead to a similar scientific conclusion when analysts make other justified analytical choices. If small, reasonable changes cause the conclusion to disappear, reverse or depend on one narrow choice, the claim may be less robust than it first appeared.
But robustness has boundaries. A result can be robust across many analyses and still come from a biased sample, a poor measurement or an unfair comparison. Robustness tests the stability of the analysis-to-conclusion link. It does not repair every earlier part of the evidence chain.
Reality Lab rule: If a conclusion depends on one fragile analysis choice, say so. If it survives several reasonable choices, confidence in that part of the evidence chain can grow.
What This Guide Owns
This guide owns one real-world evidence-transfer job: interpreting the phrase the result is robust when it refers to alternative reasonable analyses of the same data.
It does not re-own repetition, replication, statistics or scientific modelling. Those have existing owners. In particular, robustness is not the same job as:
- Reproducibility: can the reported result be recovered from the same data and analysis?
- Robustness: does the main finding survive other justified analysis choices on the same data?
- Replication: does the scientific claim survive new data collected to address the same question?
The Center for Open Science’s SCORE programme uses essentially this distinction. Reality Lab Vol No.065 already owns open data, same-evidence checking and independent new evidence. This page stays on the middle job: alternative analysis choices.
Read Reality Lab Vol No.065 on open data, reproducibility and new evidence.
The Original Reality Lab Case: The Plant-Growth Result
Imagine an original teaching dataset. Twenty similar seedlings are split into two groups. Both groups receive the same light and water. Group A receives Treatment A; Group B receives Treatment B. After two weeks, the increase in height is recorded.
The published summary says:
“Treatment A produced greater plant growth than Treatment B. The result was robust.”
What would make the word robust meaningful?
Suppose one plant in Group A grew much more than every other plant. There are several scientifically reasonable questions:
- Does Group A still look better if every valid plant is kept?
- Does the broad conclusion remain if the groups are summarised in another reasonable way?
- If the unusually large value is investigated and found to be valid, does the conclusion depend entirely on it?
- If a pre-stated rule identifies one measurement as invalid because the plant pot was knocked over, does excluding that invalid measurement change the conclusion?
If the conclusion remains similar across these justified checks, it is more robust to those choices. If the conclusion exists only when one convenient choice is made, the claim deserves a narrower wording.
Reasonable Is the Important Word
A robustness check does not mean trying random methods until one disagrees. The alternative must make scientific sense.
For example, if a question asks whether a plant grew over 14 days, comparing Day 1 with Day 14 is natural. Comparing only the minutes between 10:03 a.m. and 10:07 a.m. on Day 6 would not be a meaningful alternative for that question.
Likewise, you cannot call an analysis “robust” merely because several nearly identical calculations agree. The alternative checks should genuinely test choices that could plausibly affect the conclusion.
One Dataset Can Support Several Questions
The same table of observations may be summarised in more than one useful way. You might ask about:
- the final value;
- the amount of change;
- the rate of change;
- the typical value;
- the spread of the observations;
- how many observations cross a meaningful threshold;
- whether the pattern is similar across subgroups or time periods.
These are not interchangeable questions. A robustness check keeps the underlying scientific question stable while varying justified ways of analysing it.
Robustness Check 1: Change the Time Window Carefully
Suppose a graph shows air temperature over 30 days. A headline says, “Temperature rose strongly.” Looking only at Days 1–5 may show a steep rise. Looking at Days 1–30 may show a rise followed by a fall back toward the starting value.
If the scientific question concerns the full month, a conclusion based only on the first five days is not robust to the obvious alternative of showing the whole period.
This connects with Reality Lab Vol No.007 on selective time windows.
Robustness Check 2: Does One Unusual Result Control the Story?
An unusual observation is not automatically an error. Sometimes it is the most interesting piece of evidence. But if a headline changes completely depending on one unusual result, readers should know that the conclusion is sensitive to it.
A careful robustness check can ask two different things:
- Is the unusual result scientifically valid?
- How much does the main conclusion depend on that one result?
If it is valid, it should not simply be deleted because it is inconvenient. But the reader can still be told that the conclusion is highly sensitive to it.
Robustness Check 3: Thresholds Can Move the Count
Imagine an environmental dashboard that labels a reading “high” when it exceeds 50 units. Twelve of 30 days are classified as high. If another legitimate guideline uses 55 units, only eight days are classified as high.
The underlying measurements have not changed. The category boundary has.
A robust qualitative claim might be, “Several days were elevated under either reasonable threshold.” A more fragile claim might be, “Exactly 40% of days were high,” if that percentage exists only because one particular threshold was selected.
Robustness Check 4: Summary Choice
Suppose five fictional measurements are 4, 5, 5, 6 and 20. The value 20 is much larger than the others. A single average will be pulled upward. Another valid descriptive summary may better communicate what a typical observation looks like.
The Primary Science job is not to memorise advanced statistical rules. It is to notice when the scientific story depends heavily on one summary choice and to inspect the raw values before trusting the headline.
This is related to Reality Lab Vol No.028 on what one summary number can hide.
Robust Does Not Mean True Under Every Possible Method
Scientific communication sometimes turns a useful technical word into a magic word. “Robust” can start sounding like “unbreakable”. That is not a safe interpretation.
A result is robust only with respect to the alternatives that were meaningfully tested. A conclusion may be robust to one time-window choice but fragile to a different definition. It may survive several analysis methods while still relying on measurements from an unrepresentative sample.
Robust to what? is the question that keeps the word scientifically useful.
Robustness Cannot Repair Bad Measurement
Imagine a thermometer that is consistently 5°C too high. Five different analysts use the same biased temperature data and all reach the same conclusion. Their analyses may agree beautifully. The measurement problem remains.
This is why scientific reasoning must preserve the whole chain:
question → sampling → measurement → data → analysis → interpretation → claim.
Robustness mainly pressure-tests one region of that chain. It does not certify every link.
Worked Case 1: The Improvement Claim
A fictional product test tracks performance for 20 days. The advertisement reports a 30% improvement using Days 1–10. Using the full 20 days gives a 12% improvement.
The conclusion “there was some improvement” may survive both windows. The precise claim “30% improvement” is not robust to the longer reasonable time window.
Worked Case 2: The Unusual Plant
One plant in Treatment A grows far more than every other plant. The result is valid and the plant was not measured incorrectly. With that plant included, Group A has a much higher average. Without focusing on the average alone, most plants in the two groups look similar.
The correct response is not to secretly remove the plant. It is to report that the strong average difference depends heavily on one valid unusual observation. The finding is less robust than a simple headline suggests.
Worked Case 3: Two Reasonable Definitions
A study counts “rapid growth” as more than 3 cm per week. Another justified educational definition uses more than 2.5 cm per week. The same dataset gives different counts, but both definitions still show Treatment A producing more rapid-growth cases than Treatment B.
The exact count changes. The direction of the comparison survives. The broad conclusion is more robust than the exact percentage.
Worked Case 4: Five Analysts, One Data File
Five analysts independently receive the same anonymised dataset and scientific question. They are allowed to choose among several justified methods. Four reach the same broad conclusion; one reaches an inconclusive result.
This is more informative than forcing all five to run the identical calculation. It reveals how much the scientific conclusion depends on analysis choices. It still does not replace collecting new data.
The Robustness Audit
- Name the scientific question. Keep it stable.
- Identify the original analysis choice.
- List plausible alternative choices. They must be scientifically defensible, not random.
- Rerun the reasoning on the same data.
- Compare the conclusion, not just the exact number.
- Report sensitivity honestly. What changed? What stayed?
- Do not confuse same-data robustness with new-data replication.
- Return upstream. Check whether sampling and measurement were sound in the first place.
What Would Strengthen a “Robust” Claim?
- the alternative analyses are stated clearly;
- they are genuinely reasonable for the scientific question;
- the main conclusion survives several meaningful choices;
- the raw or underlying data are available for checking where appropriate;
- analysts do not quietly remove inconvenient valid results;
- different time windows, thresholds or summaries are justified rather than selected for a preferred outcome;
- the study also reports where the conclusion is sensitive;
- new-data studies separately test whether the claim replicates.
What Would Weaken It?
- only one analysis was tried but the result is called robust;
- alternative analyses were chosen only after seeing which ones produced the preferred story;
- small justified changes reverse the conclusion;
- the result survives only when one unusual value is excluded without a pre-stated reason;
- the word robust is used without saying what was varied;
- many analyses agree because they are almost identical;
- the analysis is stable but the measurement or sample is weak.
PSLE-Style Transfer Case
A student measures how far five toy cars travel on Surface X and Surface Y. One car on Surface X travels much farther than every other trial. The student reports only the average and says Surface X is clearly better.
What should a scientific reader do before accepting the strong conclusion?
Answer: Inspect all the individual results, check whether the unusual trial was valid, and ask whether the comparison still points in the same direction under another justified summary. Do not simply delete the unusual result or blindly trust one average.
Tempting Reasoning That Fails
- “Robust means proven.” No. It means stable to specified reasonable checks.
- “If another analysis disagrees, the original must be fraudulent.” No. Different justified methods can reveal genuine uncertainty.
- “Trying many analyses is always better.” Not if choices are arbitrary or selected only because they produce a desired answer.
- “Robustness is the same as replication.” No. Robustness can use the same data; replication adds new evidence.
- “If the conclusion is robust, the measurements must be accurate.” No. A stable analysis can still rest on weak upstream evidence.
Practice 1: Same Data, Different Window
A 30-day dataset shows an increase during Days 1–10 but almost no net change from Day 1 to Day 30. A report says, “The variable increased strongly.” What should you ask?
Answer: Ask which time window the scientific question requires and whether the conclusion survives another justified window, especially the full period.
Practice 2: Same Data, Different Threshold
A conclusion changes from “most cases are high” to “less than half are high” when a reasonable threshold moves slightly. Is the category-based claim robust?
Answer: No. The conclusion is sensitive to the threshold and should be described with that limitation.
Practice 3: Same Conclusion, Different Numbers
Three reasonable analyses estimate an improvement of 8%, 11% and 13%. All point to a modest improvement. Must robustness mean the numbers are identical?
Answer: No. Robustness often concerns whether the main conclusion remains similar, while the exact numerical estimate can vary.
Practice 4: New Data
A new team collects fresh observations and obtains the same pattern. Is that primarily a robustness check on the same dataset?
Answer: No. That is a new-data test of whether the finding replicates beyond the original dataset.
Delayed Independent Return
The next time a headline says a result is robust, do not let the word end the conversation. Make it begin one.
- Robust to which analysis choices?
- Were those alternatives reasonable?
- What stayed the same?
- What changed?
- Did the broad conclusion survive?
- Was any part of the finding fragile?
- Has the question also been tested with new data?
Teaching Guide for Parents and Tutors
Create a small invented dataset with six values: 5, 6, 6, 7, 7 and 15. Ask the learner to describe what a “typical” result looks like without teaching an advanced statistics lesson. Then ask how much the story changes if attention is placed on the unusually large value.
Next, change the question: “What if 15 is confirmed to be a valid observation rather than an error?” The learner should not delete it merely because it is inconvenient. Instead, teach the more scientific response: report the sensitivity of the conclusion.
The desired habit is not scepticism for its own sake. It is the ability to ask whether a claim rests on the data or on one fragile way of arranging the data.
Authoritative Sources
- Singapore Examinations and Assessment Board — 2026 PSLE Science Syllabus
- Ministry of Education Singapore — 2023 Primary Science Teaching and Learning Syllabus
- Center for Open Science — SCORE
- Center for Open Science — SCORE Objectives
- Center for Open Science — Large-Scale Collaboration Releases New Findings on Research Credibility
The Quiet Return
A strong scientific result should not be a story balanced on one hidden choice.
Robustness asks whether the conclusion has enough structural strength to survive other reasonable ways of reading the same evidence. When it does, that part of the claim becomes harder to knock over. When it does not, science has learned something equally useful: exactly where the conclusion is fragile.