Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

PSLE Science Reality Lab Vol No.170 | “Data Completeness = 95%” — Does That Mean the Measurements Are 95% Accurate?

PSLE-SCI-REALITY-0170

Wait, what? Ninety-five per cent complete is not ninety-five per cent correct

Imagine a science dashboard for an outdoor sensor. Beside the graph is a reassuring badge: Data completeness: 95%. A student looks at the badge and says, “Great. That means the measurements are 95% accurate.” It sounds sensible because both ideas can be written as percentages. But they answer different questions.

Completeness asks: how much of the expected data is actually present and usable? Accuracy asks a different question: how close are measurements to an appropriate reference or to the quantity we are trying to measure? A dataset can be highly complete and still contain biased measurements. It can also be incomplete yet have very good measurements during the periods that were successfully recorded.

This is a useful PSLE Science Reality Lab problem because the difficult part is not arithmetic. The difficult part is deciding what the number is evidence of.

Quick answer

If a report says 95% data completeness, the safe first interpretation is that about 95% of the observations that were expected under the stated rule were available and valid enough to count. It does not automatically mean:

  • 95% of the measurements are correct;
  • each reading is within 5% of the true value;
  • the missing 5% occurred randomly;
  • the dataset is representative of every condition;
  • the scientific conclusion is 95% certain.

The next learner job is to ask what was expected, what counted as valid, when the gaps occurred, why they occurred, and whether the missing part could change the conclusion.

The exact learner job this Reality Lab owns

This article owns one narrow real-world transfer job: reading a data-completeness percentage without turning it into a measurement-accuracy percentage. It applies existing PSLE Science skills rather than replacing them.

For the general skill of evaluating observations, information and methods, use How to Evaluate PSLE Science Observations, Information and Methods. For the difference between precision and accuracy, use How to Tell Precision From Accuracy in PSLE Science. Reality Lab Vol.170 does not re-own those skills. It shows how to deploy them when a real report puts a percentage beside the word completeness.

Rebuild the evidence object: the Orchard Roof Sensor

Consider an original composite case. A school science club places a temperature sensor on a sheltered rooftop. The system is meant to record one temperature every hour for 20 days. That means the expected number of hourly observations is:

20 days × 24 readings per day = 480 expected readings.

At the end of the project, 456 readings pass the club’s validity checks and appear in the dataset. The dashboard reports:

Routine data completeness = 95%

The calculation is straightforward: 456 ÷ 480 × 100% = 95%. But the scientific interpretation depends on what those 24 missing readings represent.

Observed

  • 480 observations were expected under the hourly schedule.
  • 456 valid observations were available.
  • 24 expected observations were not counted as valid.
  • The completeness calculation is 95% under this rule.

Claimed

“The dataset is 95% complete.”

Not yet justified

  • “The sensor is 95% accurate.”
  • “Only 5% of the readings are wrong.”
  • “The missing data do not matter.”
  • “The graph represents all weather conditions equally well.”

First check: what was the denominator?

A percentage is incomplete until you know what the whole is. Here, the denominator is 480 expected hourly observations. That is different from 480 sensors, 480 days or 480 independent experiments.

A real scientific system may define expected observations differently. A monitor might run continuously and report hourly summaries. A field station might sample once each week. A laboratory project might expect one result from each collected specimen. The completeness percentage only has meaning relative to that expected set.

This is why a learner should quietly translate the badge into a sentence:

“Ninety-five per cent of the observations expected under this reporting rule were available and valid enough to count.”

That sentence is longer than “95% complete”, but it protects the meaning.

Second check: what made a reading count as valid?

“Present” and “valid” are not always the same. A sensor may have produced a number, but the system may reject that number because the instrument was being calibrated, a quality-control check failed, a communication error corrupted the record, or the value fell outside a defined validity rule.

So when a report gives a completeness number, ask:

  • Was a reading counted only if the instrument produced any value?
  • Was a reading counted only after a quality check?
  • Were estimated or substituted values counted?
  • Were periods of planned maintenance included in the expected total?
  • Was completeness calculated for the whole year, one season or only the period when the monitor was operating?

These details do not automatically make a dataset good or bad. They define what the completeness label means.

Third check: where are the missing readings?

Two datasets can both be 95% complete and still have very different scientific usefulness.

DatasetAvailable readingsMissing patternPossible consequence
A456 of 48024 isolated hours scattered across many ordinary daysMay leave most conditions reasonably represented
B456 of 480All 24 missing hours occurred during the hottest dayCould hide the very event used to judge extreme heat

The completeness percentage is identical. The evidence pattern is not.

This gives a powerful Reality Lab habit: when data are missing, location in time or condition can matter as much as the amount missing.

Missing at random is not something you should assume

Suppose the rooftop sensor works well in mild conditions but shuts down whenever direct afternoon sun heats the electronics above their operating range. The missing readings would then occur more often during the hottest conditions. If a student simply ignores the gaps, the remaining data could make the roof look cooler than it really was during the study period.

Now suppose instead that a network cable disconnects for one random hour on several different days, unrelated to temperature. The missing pattern may have a smaller effect on the temperature story.

Neither pattern can be diagnosed from “95% complete” alone. The badge is a starting point, not the whole investigation.

Completeness, accuracy, precision and representativeness are different jobs

QuestionEvidence idea
How much expected data is available?Completeness
How close is a measurement to an appropriate reference or target?Accuracy / bias evaluation
How closely do repeated measurements agree?Precision / repeatability
Does the collected evidence cover the conditions or population the claim is about?Representativeness

One dataset can score well on one dimension and poorly on another. For example, a miscalibrated sensor can produce a beautifully complete dataset: every expected hour is present, but every value is shifted upward. Completeness is excellent; accuracy is poor. Conversely, a well-calibrated instrument can make excellent measurements during the hours it operates but lose power for several days. Accuracy may be good; completeness is poor.

Worked case 1: the perfect-looking 100%

A product-comparison test records drying time for two materials every minute for one hour. All 60 expected readings appear for both materials, so completeness is 100%. Later, the investigator discovers that the timer was running 10% fast.

Tempting reasoning: “The data are 100% complete, so the measurements are reliable.”

Better reasoning: All expected readings are present, so completeness is 100%. However, completeness does not test whether the time measurements are accurate. The timer problem is a separate measurement-quality issue.

Worked case 2: lower completeness, stronger measurement quality

Two thermometers are checked against an appropriate reference before use. During a 100-hour investigation, one instrument records only 88 valid hours because its battery fails. The available measurements agree well with the reference checks before and after the study.

The dataset is only 88% complete under the hourly schedule. That does not mean the 88 available measurements are 88% accurate. It means 12 expected hours are missing. To judge measurement accuracy, we need the reference evidence. To judge the conclusion, we also need to know which 12 hours are absent.

Worked case 3: the missing event problem

A river-level logger is expected to record every 15 minutes for four weeks. Its overall completeness is 97%. Almost all missing observations occurred during one severe storm because debris damaged the sensor.

If the question is “What was the ordinary river level during most of the month?”, the remaining data may still provide useful evidence. If the question is “What was the highest level during the storm?”, the missing 3% could be crucial. The value of completeness depends on the scientific question.

Representation check: a green badge can make a limited metric feel like a universal quality score

Dashboards often use colour: green for high completeness, amber for moderate completeness, red for poor completeness. Colour can be useful, but it can also encourage a shortcut: green = good data in every sense.

Do not let the badge do more scientific work than its definition allows. A green completeness indicator may support the statement “few expected observations are missing under this rule.” It does not automatically certify calibration, representativeness, absence of bias, correct processing or correct interpretation.

Comparison and baseline check

If two monitoring sites report 98% and 93% completeness, it is tempting to rank the first as scientifically “better”. Before doing that, check whether the completeness calculations use the same expected schedule and validity rules. One site might be expected to report every minute, another every hour. One system might exclude planned maintenance from the denominator, another might include it. The percentages are comparable only when the measurement job and counting rule are comparable.

Method and variable check

For a completeness claim, the most useful method questions are often operational rather than conceptual:

  • What schedule created the expected count?
  • What caused an observation to be rejected?
  • Was equipment failure related to the measured condition?
  • Were missing readings concentrated at one site, one time or one extreme condition?
  • Were estimated replacements inserted, and are they distinguished from direct measurements?
  • Did the reporting period begin or end when the instrument was inactive?

Alternative explanations for a low completeness number

A low completeness percentage does not always mean careless science. Possible explanations include instrument failure, power loss, blocked communication, planned maintenance, unsafe field conditions, failed quality-control checks, a damaged sample, or a deliberate decision not to invent a value when no defensible observation exists.

Sometimes not filling a gap is more scientifically honest than pretending to know what happened.

What evidence would strengthen the claim?

  • A clear definition of the expected observations.
  • A stated rule for what counts as valid.
  • A timeline showing exactly where gaps occurred.
  • A reason for missing observations where known.
  • Separate calibration or reference checks for measurement accuracy.
  • An analysis of whether missing periods differ systematically from observed periods.
  • Transparent flags showing measured, estimated and missing values.

What evidence would weaken a broad claim?

  • Missing data concentrated during extreme conditions.
  • A completeness calculation whose denominator is unclear.
  • Changing validity rules midway through the record without explanation.
  • A dashboard that silently fills gaps but still labels the result as measured data.
  • Using a completeness percentage as if it directly measured accuracy.

How far can the conclusion travel?

Suppose a station reports 95% completeness for August. That supports a claim about the amount of expected August data available under the stated reporting rule. It does not automatically support completeness in July, during another year, at another station or under another sampling schedule.

This is a general scientific habit: the conclusion should travel only as far as the evidence travels.

PSLE-style transfer case

An original investigation records light intensity every hour for 10 days. It should contain 240 readings. Only 216 readings are valid, so the dataset is 90% complete. All 24 missing readings occurred from 12 noon to 2 pm on several sunny days.

A student concludes: “The dataset proves the average light intensity was low because 90% of the readings were complete.”

A stronger answer would say that 90% completeness only describes how much expected data is available. Because the missing observations are concentrated around bright midday periods, the missing pattern could lower the calculated average from the remaining readings. The completeness percentage alone therefore cannot prove that the average light intensity was low.

Tempting but invalid reasoning

  • “95% complete means 95% correct.” Wrong quantity.
  • “Only 5% is missing, so it cannot matter.” A small missing fraction can contain the event that matters most.
  • “100% complete means perfect data.” Completeness does not test calibration, bias or interpretation.
  • “Lower completeness means the study is useless.” It may still answer some questions well if the missingness is understood and the conclusion is limited appropriately.
  • “We can just fill every gap.” Estimated values must not be silently converted into observations.

Measurement and model limits

A completeness metric is itself a model of the record. It requires a definition of expected observations and valid observations. Change those definitions and the percentage can change. That does not make the metric meaningless; it means the definition is part of the evidence.

For younger scientists, this is an important lesson: a number can be mathematically correct and still be scientifically misunderstood if we attach it to the wrong question.

Independent return: try these without looking back

  1. A sensor expected 1,000 readings and returned 990 valid readings. What does 99% completeness tell you? What does it not tell you?
  2. Two stations are each 95% complete. At Station A the gaps are scattered randomly. At Station B every gap occurs during heavy rain. Which additional question matters before comparing rainfall?
  3. A lab report is 100% complete but a reference check shows every measurement is shifted upward. Which quality idea is affected?
  4. A dashboard has no blank spaces because software estimated every missing hour. What distinction should the reader ask the provider to show?

Explained answers

1. It tells you that 990 of 1,000 expected readings counted as valid under the stated rule. It does not tell you measurement accuracy, where the missing ten occurred, or whether the record is representative.

2. Ask whether the missing heavy-rain periods could remove large rainfall values from Station B. Same percentage, different evidence pattern.

3. Measurement accuracy or systematic bias is the problem, not completeness.

4. Ask which values were directly measured and which were estimated or filled. A seamless record can hide provenance differences.

For parents and tutors: teach the noun after the percentage

When a child sees a percentage, ask them to say the noun with it: 95% of what? Then ask the scientific job: What question does this percentage answer?

Useful prompts include:

  • “What was expected?”
  • “What counted as present?”
  • “What is missing?”
  • “Where is it missing?”
  • “Does this number describe quantity of data, quality of measurement, or something else?”

Resist giving a universal template. The aim is for the learner to identify the evidence object, not to memorise a sentence about percentages.

Why this belongs in PSLE Science

The 2026 PSLE Science assessment objectives include interpreting and analysing information, evaluating observations, information and methods, and communicating explanations and reasoning. The 2023 Primary Science syllabus also treats scientific inquiry as more than recalling facts: learners examine evidence, assumptions, uncertainty and how scientific information is communicated. A completeness badge is therefore a useful real-world transfer object because it asks the learner to connect a familiar percentage to the exact scientific meaning it carries.

Authoritative sources

Quiet return

A scientific number becomes useful when you know the job it is doing. 95% complete can be excellent evidence about how much expected data is available. It is not a shortcut to 95% accurate. Keep the label attached to the right question, inspect the missing part, and let the evidence say exactly as much as it can.