Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

PSLE Science Reality Lab Vol No.257 | “Detected in 70% of Samples” — Was It Present 70% of the Time or Across 70% of the Area?

Stable internal ID: PSLE-SCI-REALITY-0257

Wait, what? A monitoring report says, “Chemical M was detected in 70% of samples.” A headline rewrites that as, “Chemical M is present 70% of the time.” An infographic goes further: “70% of the river is contaminated.”

Those three sentences sound as if they carry the same percentage. They do not carry the same scientific object.

“Detected in 70% of samples” has a denominator: the samples that were collected and analysed under a stated sampling design and detection rule. It does not automatically tell us 70% of hours, 70% of river length, 70% of sites, 70% of animals or 70% of a population. To move from the sample percentage to one of those wider claims, we need evidence showing that the samples represent that wider target.

This Reality Lab owns one narrow learner job: how to read a detection-frequency percentage by reconstructing what was sampled, how often, where, with what detection threshold and with what sampling effort before deciding what the percentage can represent.

Quick Answer

No. If 70 of 100 analysed samples contain a reportable detection, then the observed sample detection frequency is 70%. That is a statement about those samples. Whether it also represents time, area, sites or a wider population depends on how the samples were selected and how detection was defined.

  • Count the samples in the denominator.
  • Check whether several samples came from the same place or time.
  • Check whether sampling effort was equal across locations and periods.
  • Check the reporting or detection threshold.
  • Keep “not detected” separate from “proved absent.”
  • Do not silently replace “percentage of samples with detections” with “percentage of the world where the thing exists.”

The Exact Learner Job This Page Owns

This page is not a generic owner of percentages, probability, sampling, ecology, environmental chemistry or statistics. Existing eduKateSengkang guides remain the owners of those foundations. Reality Lab Vol No.257 applies them to a common real-world communication object: a detection-frequency statement in a monitoring report, map, infographic, news story or scientific summary.

It is also distinct from Reality Lab Vol No.243 on species occupancy. Occupancy can be a model-based estimate of which sites are occupied while accounting for imperfect detection. Here the job is earlier and simpler: what does the raw percentage of samples with detections actually describe?

Start With the Denominator

Suppose a report says:

Substance R was detected in 70% of samples.

Your first question should not be, “Is 70% a lot?” Your first question should be, “70% of what?”

If there were 100 samples and 70 had reportable detections, the arithmetic is straightforward. But the scientific meaning depends on the architecture of those 100 samples.

Sampling designWhat 70/100 most directly describes
100 samples from one station over timeDetection frequency among those sampled times at that station.
One sample from each of 100 stations on one dayDetection frequency among those sampled stations on that day.
Ten samples from each of ten stationsDetection frequency among 100 sample events, with repeated observations within stations.
Seventy samples from Site A and thirty from all other sitesA sample-weighted frequency dominated by Site A, not an equal-area summary.

The same fraction—70/100—can sit on four very different scientific designs.

Rebuild the Object: An Original River Survey

Imagine River Cedar has four monitoring stations: North, Central, South and Estuary. Scientists collect samples for Chemical K.

StationSamples collectedDetections
North102
Central106
South1010
Estuary7052
Total10070

The report can truthfully say, “Chemical K was detected in 70% of the 100 samples.” But can you say “70% of the four stations contained Chemical K”? No. All four stations had at least one detection, so the station-level answer would be 4 of 4 for this particular dataset.

Can you say “70% of the river length contained Chemical K”? No. The stations do not divide the river into equal pieces, and the Estuary received seven times as many samples as each other station.

Can you say “Chemical K was present 70% of the year”? No. We have not yet been told when the samples were collected or how evenly they cover time.

This is the core habit: preserve the denominator’s identity.

Observed, Claimed and Inferred

LayerRiver Cedar example
Observed70 of 100 collected samples produced reportable detections.
Direct claimObserved sample detection frequency = 70%.
Possible wider claimThe target is commonly encountered under the monitored conditions.
Unsupported leap without more design evidence70% of river area, 70% of days, or 70% of water contains the target.

USGS Uses “Detection Frequency” Carefully

U.S. Geological Survey monitoring publications report detection frequency as how often a compound was detected in samples. USGS guidance also shows why the detection threshold matters: compounds can have different reporting capabilities, so detection frequencies may need to be calculated at a common concentration threshold before fair comparisons are made.

That point is powerful for Primary Science. “Detected” is not a universal yes/no property floating independently of the method. A detection statement belongs to a method, sample and reporting rule.

Detection Threshold: The Hidden Rule Behind the Percentage

Suppose Laboratory A can reliably report Chemical K down to 0.01 µg/L, while Laboratory B reports only values at or above 0.10 µg/L. The same water samples could produce a higher detection frequency under A’s more sensitive reporting rule.

If someone compares “80% detected” from A with “40% detected” from B without checking thresholds, they may think the environments differ more than they really do.

USGS explicitly warns that detection frequencies for different compounds should not be directly compared when reporting levels differ, and it uses common thresholds to make comparisons more meaningful.

Worked Case 1: 70% of Samples Is Not 70% of Time

A sensor station is sampled ten times on Monday, ten times on Tuesday, and eighty times during one storm on Wednesday. Seventy of the hundred samples show Chemical P, mostly during the storm.

The sample detection frequency is 70%. But the samples do not represent one hundred equally spaced moments. Wednesday was sampled far more heavily. It is therefore unsafe to say “Chemical P was present for 70% of the three days” unless the sampling design supports that time weighting.

Worked Case 2: 70% of Samples Is Not 70% of Area

A lake survey collects ninety samples near an inlet and ten samples from the rest of the lake. Seventy inlet samples are positive; none elsewhere are positive. Overall detection frequency is 70%.

Does that mean 70% of lake area has the target? No. The survey deliberately concentrated sampling where the target was expected. That may be a sensible monitoring strategy, but sample frequency cannot be turned into area coverage unless sampling locations represent area appropriately.

Worked Case 3: Many Samples From One Site

Site A has 90 samples and 63 detections. Site B has 10 samples and 7 detections. Overall detection frequency is 70%. A headline says “70% of sites were positive.”

There are only two sites, and both had detections. The site-level fraction is therefore 2 of 2, not 70%. The headline changed the unit of the denominator from samples to sites.

This is one of the most important evidence-reading habits in the whole Reality Lab series: never let a denominator change its identity while keeping the same percentage.

Worked Case 4: Non-Detect Does Not Automatically Mean Absent

Thirty samples do not produce reportable detections. A student says, “So Chemical K definitely was not present in those thirty samples.”

That conclusion is too strong. A non-detect can mean the target was absent, but it can also mean any amount present was below the method’s detection or reporting capability, or that sample and method conditions limited detection. The exact interpretation depends on the method.

This article does not re-teach the detection-limit owner. It uses that existing idea for one purpose: the 70%/30% split should not be treated as 70% certainly present and 30% certainly absent unless the evidence supports that stronger statement.

Worked Case 5: Equal Samples, Unequal Effort

Two wildlife teams each collect 50 environmental samples. Team A searches for eight hours per site before collecting. Team B samples the first convenient location after ten minutes. Team A reports detections in 60% of samples, Team B in 30%.

Can the percentages be compared directly as if only the environment differed? Not safely. Sampling effort and placement differ. The two sets of samples may have different chances of encountering the target.

Worked Case 6: The Denominator Changed Between Years

In Year 1, scientists sample 20 locations once each and get 10 detections: 50%. In Year 2, they sample the five historically positive locations ten times each and get 35 detections from 50 samples: 70%.

Can a graph label this “detection increased from 50% to 70%” and imply the river became more affected? The arithmetic percentages are real, but the sampling design changed. The increase could partly reflect where and how often scientists looked.

Before reading a trend, demand a comparable denominator.

Worked Case 7: Frequency Is Not Concentration

Substance A is detected in 90% of samples, usually just above the reporting threshold. Substance B is detected in 20% of samples, but some detections are much higher in concentration.

Which substance has “more pollution”? The detection-frequency percentages alone cannot answer. Frequency of occurrence and concentration magnitude are different evidence dimensions. Depending on the scientific question, both may matter.

Again, this page does not turn into risk assessment. It simply protects the measurement object: how often detected is not how much detected.

The Four Denominator Questions

Whenever you see “detected in X% of samples,” ask four questions in order:

  1. How many samples? A percentage without a count can hide a tiny evidence base.
  2. Samples of what? Water bottles, air filters, soil cores, swabs, camera events and biological specimens are not interchangeable.
  3. Collected where and when? The sampling frame controls what the percentage can represent.
  4. Detected by which rule? Method sensitivity and reporting threshold affect the detection count.

Only after those four questions should you ask whether the percentage can generalise to a wider target.

Representation Check: Maps Can Make the Denominator Disappear

Suppose an infographic colours an entire district orange and writes “70% detection.” The colour fills space, so the viewer may assume 70% of the land area is affected. But perhaps the original study collected 100 water samples at only twelve monitoring sites.

The map has converted a sample statistic into a spatial-looking object. That can be useful if the mapping method is explained. It becomes misleading if the filled area visually implies coverage that the sampling design did not establish.

Comparison Check: Use the Same Detection Basis

USGS examples show why scientists sometimes calculate detection frequencies at common concentration thresholds. Imagine:

  • Chemical A can be reported down to 0.01 µg/L.
  • Chemical B can be reported only down to 0.10 µg/L.
  • A is “detected” more often partly because the method sees smaller concentrations.

If the scientific question is “Which chemical occurs more frequently above 0.10 µg/L?”, recalculate both at that shared threshold. Fair comparisons require a shared rule.

Method and Variable Check

Detection frequency can change even when the underlying environment does not change, if the monitoring design changes. Check:

  • number of samples;
  • sampling locations;
  • sampling dates and seasons;
  • sample depth or material;
  • sample volume;
  • laboratory method;
  • detection/reporting threshold;
  • sample handling and storage;
  • targeted versus random or systematic selection.

You do not need all factors to be identical in every real study. You need enough information to know which differences matter to the claim.

Alternative Explanations for a Higher Detection Frequency

If detection frequency rises from 40% to 70%, one explanation is that the target truly became more common in the sampled environment. But other explanations include:

  • a more sensitive method was introduced;
  • sampling moved closer to likely sources;
  • more samples were collected during high-flow or high-activity periods;
  • the reporting threshold changed;
  • sample volume increased;
  • sample preservation improved;
  • the sample mix changed from broad surveillance to targeted investigation.

A good report tries to separate these possibilities. A good reader at least asks whether they were considered.

What Evidence Would Strengthen a Wider Claim?

  • A sampling design explicitly built to represent the target area, time period or population.
  • Clear counts of samples and detections.
  • Sampling effort distributed appropriately across the target.
  • A stable or harmonised detection threshold for comparisons.
  • Repeated sampling across relevant seasons or conditions.
  • Transparent treatment of non-detects and missing samples.
  • Independent checks showing that the sampled units resemble the wider target.
  • Language that distinguishes “sample detection frequency” from “estimated prevalence” or “area affected.”

What Would Weaken It?

  • The denominator is missing.
  • Many samples come from a small number of repeatedly sampled sites.
  • Sampling concentrates on suspected hotspots but the result is generalized to the whole area.
  • Detection limits differ between groups without adjustment.
  • Sampling periods differ strongly between comparisons.
  • “Not detected” is treated as certain absence.
  • Frequency of detection is mistaken for amount, severity or risk.
  • A percentage of samples is redrawn as a percentage of map area without a spatial model or representative design.

How Far Can the Conclusion Travel?

The strongest immediate statement from “70 of 100 samples detected” is exactly that: 70 of these 100 analysed samples met the study’s detection rule.

With a strong design, scientists may estimate wider occurrence. But the route depends on what the samples represent. Do not leap automatically from samples to:

  • percent of time;
  • percent of area;
  • percent of sites;
  • percent of individuals;
  • percent of total material;
  • or probability that a new randomly chosen unit is positive.

Each target has its own denominator.

Tempting but Invalid Reasoning

  • “70% of samples positive = 70% of the river positive.” Samples are not square kilometres.
  • “70% positive = present 70% of the time.” Sample events may not be evenly distributed across time.
  • “30% non-detect = absent in 30%.” Non-detects are bounded by method capability.
  • “Higher detection frequency = higher concentration.” Frequency and magnitude are different.
  • “More samples automatically means more representative.” One thousand biased samples can still misrepresent a target.
  • “Same percentage means same evidence.” 7/10 and 700/1000 have the same percentage but different resolution and sampling histories.

PSLE-Style Transfer Case: The Pond Survey

Original classroom case: A school science team collects 20 water samples from three ponds. Pond A contributes 2 samples, Pond B contributes 3, and Pond C contributes 15 because it is beside the school. Fourteen of the 20 samples show Marker Q, so the team writes “Marker Q detected in 70% of samples.”

A display board changes this to: “70% of the ponds contain Marker Q.”

What is wrong?

  • The original denominator is samples, not ponds.
  • Pond C dominates the sample count, so the 70% is heavily weighted toward one pond.
  • To claim a percentage of ponds, each pond’s status needs to be defined and the pond denominator used.
  • If all three ponds have at least one detection, the observed pond-level statement is 3 of 3 sampled ponds, not 70%.
  • Neither statistic automatically describes every future day because time was not part of the denominator.

Transfer Case: Comparing Two Years

Year 1: 20 evenly spaced monthly-and-site samples, 8 detections. Year 2: 100 samples concentrated after storms, 60 detections. A headline says detection frequency rose from 40% to 60%.

The numbers are correct for the sample sets. But the sampling schedule changed strongly. A learner should ask whether a comparable subset or common design is needed before interpreting the difference as environmental change.

Explained Practice

Practice 1 — Identify the Denominator

“Detected in 18 of 30 samples.” What is the direct detection frequency?

Answer: 60% of the analysed samples. Do not rename the denominator.

Practice 2 — Sites or Samples?

Thirty samples came from five sites. Twenty samples came from Site A. Can 60% sample detection be reported as “60% of sites”?

Answer: No. Site-level status needs a separate calculation using the five sites.

Practice 3 — Time?

Eighty of 100 samples were collected during two storms. Can 80% sample detection tell you the target was present 80% of the month?

Answer: No. Sampling was concentrated in particular events, so sample frequency is not month-time coverage.

Practice 4 — Detection Threshold

Method A detects down to a lower concentration than Method B. Why can raw detection percentages be unfair to compare?

Answer: A can count low-level detections that B cannot report. Use a common threshold or otherwise account for method differences.

Practice 5 — Frequency Versus Amount

Substance X is detected in 90% of samples and Y in 20%. Can you conclude X always has the higher concentration?

Answer: No. Detection frequency does not tell you the concentration magnitude in each positive sample.

Delayed Independent Return

On a later day, give the learner five statements:

  • “positive in 7 of 10 samples”;
  • “positive at 7 of 10 sites”;
  • “positive on 7 of 10 days”;
  • “70% of the mapped area”;
  • “70% of animals tested.”

All can equal 70%, but ask the learner to draw a box around each denominator and explain what evidence would be needed to convert one statement into another. This delayed return is a strong test of whether denominator identity has become a real reasoning habit.

Parent and Tutor Teaching Guide

Use coloured counters. Put 100 counters in a bag, but say seventy counters were sampled from one corner and thirty from everywhere else. Ask whether “70% of samples” automatically describes “70% of the bag.” The physical mismatch makes the denominator problem visible.

Then move to time. Place ten sample cards on a calendar, with eight cards on one rainy day and two spread across the rest of the month. Ask whether sample percentages can be read directly as time percentages.

The teaching goal is not to make the child suspicious of every percentage. It is to make the child ask, calmly and automatically: what exactly was counted in the bottom of this fraction?

Routes to Existing Canonical PSLE Science Owners

Authoritative Sources

Quiet Return

A percentage can look complete because it ends with a percent sign. Scientifically, it is incomplete until you know the denominator.

When you read “detected in 70% of samples,” keep the word samples attached to 70%. Then inspect where, when and how those samples were collected, and what counted as a detection. Only after that may the conclusion travel.

Do not let a percentage keep its number while secretly changing what it counts.