PSLE-SCI-REALITY-0065
Wait, What? A study can share every data file and still not have been independently repeated.
You open a scientific article and see a reassuring message: Data available here. There is a download button. The spreadsheet is public. The methods are described. It feels as if the result has already passed a powerful test.
Something important has happened. Other people can inspect the evidence more directly. They may be able to check calculations, look for missing values, try another analysis and see whether the reported numbers can be recovered from the shared files.
But downloading the same dataset is not the same as carrying out the investigation again with new observations.
Open data improves transparency and checkability. It does not magically create a second independent experiment.
Quick Answer
“Open data available” means the underlying data have been made accessible under some stated conditions. That can make a scientific claim easier to inspect and reanalyse. It does not by itself show that another team has collected new evidence and obtained a similar result. To evaluate the claim, distinguish four jobs: accessing the data, checking the analysis, reproducing the reported result from the same evidence, and testing the scientific question again with independent evidence.
Reality Lab rule: Shared evidence is easier to check. It is not automatically new evidence.
What This Guide Teaches—and What It Does Not
This Reality Lab owns one transfer job: how a Primary 5/6 learner should interpret an “open data”, “data available” or “download the dataset” statement in scientific communication.
It does not teach the full professional vocabulary of open science. Different scientific fields sometimes use the words reproducibility and replicability differently. The learner job here is more durable than the terminology: ask whether someone is checking the same evidence or collecting new evidence.
eduKate already owns the broader PSLE distinction between repeatability and reproducibility. This page applies that reasoning to a public data-sharing claim rather than re-teaching the micro-skill.
The Original Reality Lab Case: The Water-Filter Spreadsheet
Imagine an original teaching case. A student research team tests two water-filter materials using 40 prepared water samples. Their article reports:
“Filter A removed more suspended particles than Filter B. All data are openly available.”
The shared folder contains a spreadsheet with 40 before-and-after measurements and a second file that calculates the group averages.
Another student downloads the files, checks every formula and obtains exactly the same averages. What has been shown?
- The reported averages can be recovered from the shared dataset.
- The arithmetic in the shared calculation appears consistent.
- The data are available for further checking.
What has not yet been shown?
- That the measurements were recorded correctly in the first place.
- That the experiment was free of uncontrolled differences.
- That a new batch of filters and water samples would produce the same pattern.
- That another independent team would reach the same conclusion.
The same evidence has been checked more deeply. The world has not yet supplied a second set of evidence.
Four Different Scientific Jobs
| Job | What happens | What it can add |
|---|---|---|
| Data access | Another person can obtain the shared data. | Transparency and inspectability. |
| Reanalysis | Another person applies calculations or alternative analyses to the same data. | Checks arithmetic, coding, assumptions and sensitivity. |
| Same-evidence reproduction | Another person follows the documented analysis and recovers the reported result from the same dataset. | Confidence that the result follows from the available data and procedure. |
| Independent new-evidence test | A new investigation collects new observations or samples. | Tests whether the scientific pattern survives beyond the original dataset. |
These jobs support one another, but they are not interchangeable.
Why Open Data Is Scientifically Valuable
Major scientific organisations encourage appropriate data sharing because other researchers can inspect evidence, verify analyses, combine compatible datasets, identify errors and build new research from earlier work. The Royal Society requires data and code needed to reproduce reported results to be made available when appropriate, subject to ethical and practical limits.
The Center for Open Science similarly promotes practices that make research more transparent and reusable.
For a Primary Science learner, the important idea is simple: science becomes easier to challenge when the evidence trail is visible.
But “Available” Is Not the Same as “Understandable”
A folder containing 20 mysterious files is technically accessible but may still be difficult to check. Useful shared data normally need context.
- What does each column mean?
- What units were used?
- Which values are missing?
- Which rows were excluded, and why?
- Were values raw, corrected or calculated?
- Which version of the dataset produced the paper?
- What processing steps happened before the shared file was created?
Without that information, the data can be open but scientifically opaque.
Open Data Cannot Repair a Weak Investigation by Itself
Suppose two plant groups received different fertilisers, but one group also received more sunlight. The full dataset is shared publicly. Does openness turn the comparison into a fair test?
No. Transparency helps us see the flaw. It does not remove the flaw.
This is an important distinction. Open science can make weaknesses easier to discover, but evidence quality still depends on the design, measurement and reasoning that produced the data.
Same Data, Different Analysis
Two researchers can analyse the same dataset in different reasonable ways. They may choose different time windows, definitions, exclusions or models. If the conclusion changes dramatically under small reasonable choices, that is scientifically important.
Open data allow those choices to be pressure-tested. A reader can ask:
- Does the pattern remain if one unusual result is kept?
- Does it remain if a different baseline is used?
- Does it remain if the full time period is shown?
- Does it remain if groups are compared in another justified way?
This kind of reanalysis can strengthen or weaken confidence. It still uses the original evidence.
New Data Ask a Different Question
Now imagine a second school buys new samples of Filter A and Filter B, writes its own test plan, carries out 40 fresh trials and again finds that Filter A removes more suspended particles.
This adds something the download-and-recalculate exercise could not: new observations from a separate run.
If the second team changes important conditions as well, the new work may also test whether the finding extends beyond the original set-up. eduKate’s existing follow-up investigation guides own the precise distinction between repeating and extending an investigation.
A Public Data File Can Still Be Incomplete
“Data available” is a claim that deserves its own audit. Ask whether the shared material includes the data needed for the published result, not merely a selected summary.
A dataset may omit:
- invalid trials that were removed under a stated rule;
- personal or sensitive information that cannot ethically be shared;
- intermediate processing files;
- proprietary measurements;
- data that were never preserved;
- variables needed to reconstruct the exact published analysis.
Some restrictions can be legitimate. The scientific question is whether the remaining evidence and documentation are sufficient for the claim being made.
Version Matters
Datasets can be corrected after publication. A later download may not be identical to the file used in the original analysis. Strong data repositories therefore use version numbers, dates, identifiers or change histories.
This connects to Reality Lab Vol No.026 on scientific corrections and Vol No.064 on adjusted data. A reproducible evidence trail needs to identify which version was actually used.
The Open-Data Audit
- What exactly is shared? Raw observations, processed data, summary table, code, spreadsheet or only selected results?
- Is the dataset documented? Units, variables, missing values, exclusions and processing steps should be interpretable.
- Can the reported result be recovered from the shared evidence?
- Can another reasonable analysis be tried?
- Has anyone collected new independent evidence?
- Which version of the data was used?
- Are there legitimate privacy, ethical or legal limits on sharing?
Worked Case 1: Same Spreadsheet, Same Answer
A second student downloads a public spreadsheet, repeats the average calculation and gets exactly the number in the article.
This supports the claim that the reported average follows from the shared values. It does not independently confirm that every original measurement was accurate or that a new experiment would produce the same pattern.
Worked Case 2: The Hidden Exclusion Rule
A dataset contains 100 rows, but the published graph uses only 82. The article does not explain which 18 were omitted.
Open data have made an important question visible. The next step is to identify the exclusion rule and whether it was scientifically justified. The mere existence of the full file does not make the published analysis correct.
Worked Case 3: New Team, New Samples
A new team follows the published method using new samples and obtains a similar result.
That provides a different kind of support from simply reanalysing the original file. The scientific conclusion has now survived contact with new evidence.
Worked Case 4: Open Code, Missing Instrument Settings
A project shares all its computer code and processed data, but the method does not record an important instrument setting used during collection.
The analysis may be reproducible from the processed file, while the physical measurement is difficult to repeat faithfully. Transparency has improved one part of the evidence chain but not the whole chain.
What Would Strengthen an Open-Data Claim?
- complete data needed for the reported analyses;
- clear units and variable definitions;
- documented exclusions and missing values;
- analysis code or formulas;
- version identifiers and change history;
- a detailed method linking observations to the files;
- independent reanalysis;
- new-data follow-up studies.
What Would Weaken It?
- a download link that does not contain the evidence used in the paper;
- unexplained missing rows or columns;
- undocumented processing;
- no way to identify the dataset version;
- a headline that calls reanalysis “independent replication”;
- the shared file contains only already-summarised numbers when the claim depends on underlying observations.
PSLE Science Transfer: Copying the Same Table Is Not a New Trial
A class carries out an investigation and records five results. Student A calculates the average. Student B uses the same five numbers and independently calculates the same average.
Student B has checked the calculation. The class still has only five experimental results.
To add new experimental evidence, the investigation must produce additional valid observations. That simple distinction is the bridge from classroom science to open scientific data.
Tempting Reasoning That Fails
- “The data are public, so the conclusion must be true.” Openness improves checkability, not certainty.
- “Someone downloaded the data and got the same answer, so the experiment was independently repeated.” The same evidence was reanalysed.
- “If data are not fully public, the research must be dishonest.” Privacy, ethics, law and legitimate restrictions can limit sharing.
- “Open data mean raw data.” A shared dataset can be processed or summarised.
- “More files mean stronger evidence.” Relevance, documentation and provenance matter more than file count.
Practice 1: The Download Button
A news article says, “The study is fully verified because anyone can download its data.” What is wrong with the sentence?
Answer: Download access improves transparency but does not guarantee that the data were collected correctly, the analysis was appropriate or new independent evidence agrees.
Practice 2: Same Data, Different Result
Two analysts use the same public dataset. One includes all valid results; the other excludes a group under a different justified rule. Their summaries differ. What should happen next?
Answer: Compare the rules and test how sensitive the conclusion is to those choices. The disagreement can reveal an important analysis boundary.
Practice 3: New Evidence
Which adds more independent evidence: ten people recalculating the same 40 measurements, or one new team collecting 40 fresh valid measurements under the same method?
Answer: The ten recalculations can strongly check the analysis, but the new team contributes new observational evidence about whether the pattern survives another investigation.
Delayed Independent Return
The next time a scientific page says “data available”, ask two separate questions:
- What can another person check using these same data?
- Has anyone tested the scientific claim using new evidence?
If you keep those questions separate, the phrase “open data” becomes informative instead of magical.
Teaching Guide for Parents and Tutors
Give the learner five invented measurements and ask them to calculate an average. Then calculate the same average yourself. Ask: “How many measurements do we have now?” The answer is still five, not ten.
Next, collect five new measurements from a second simple, safe classroom observation. Now ask what changed. The learner should see the distinction between checking the same evidence and adding new evidence.
The weak link to watch is the assumption that transparency, reproducibility and independent replication are all the same property. Keep the jobs separate before adding professional terminology.
Authoritative Sources
- Singapore Examinations and Assessment Board — 2026 PSLE Science Syllabus
- Ministry of Education Singapore — 2023 Primary Science Teaching and Learning Syllabus
- The Royal Society — Open Science
- The Royal Society — Data Sharing and Mining
- Center for Open Science — Transparency and Open Research Infrastructure
The Quiet Return
Open data do something important: they let more people look inside the evidence chain.
But looking again at the same evidence and asking reality for new evidence are different scientific acts. Strong science needs both.