Series ID: PSLE-SCI-REALITY-0282
Wait, What? Sixty Percent of the DNA Reads Is Not Automatically Sixty Percent of the Animals
A fictional pond report shows a colourful bar chart. Species A has 60% of the environmental-DNA sequencing reads, Species B has 25%, and Species C has 15%. A student immediately writes: “Therefore 60% of the animals in the pond were Species A.”
The chart looks like a population pie. But it is not necessarily one. Environmental DNA, or eDNA, can be collected from water, soil or other environmental samples. In metabarcoding, DNA is extracted, selected genetic regions are amplified and many sequence reads are generated and assigned to taxa. The number of reads at the end of that pipeline is affected by biology and by the sampling and laboratory process.
U.S. Geological Survey research provides a useful real-world warning. Read numbers can sometimes contain information related to abundance, but taxon-specific amplification and sequencing biases can distort community composition, and in one USGS-supported study the relationship was stronger for dominant taxa and weaker for non-dominant species. So the scientifically careful conclusion is neither “read counts are meaningless” nor “read share equals animal share.” The correct question is: what has been validated for this method, taxon, sample and purpose?
Quick Answer
- eDNA sequencing reads are DNA-sequence observations produced by a sampling and laboratory pipeline.
- A species receiving 60% of reads does not automatically mean it formed 60% of the living animals, individuals or biomass.
- Organisms can shed different amounts of DNA.
- Sampling can collect DNA unevenly across places and times.
- DNA extraction can recover some material more efficiently than other material.
- Primers and PCR can amplify taxa differently.
- Gene-copy differences and sequencing processes can affect read proportions.
- Reference databases and bioinformatic rules affect which reads are assigned to which taxa.
- Read abundance may support bounded abundance inferences in some validated systems, but that relationship must be demonstrated rather than assumed.
- The student should keep detection, read abundance, organism abundance and biomass as separate scientific quantities.
The Exact Learner Job This Reality Lab Owns
This volume owns one evidence-transfer job: how a Primary 5/6 learner should evaluate an eDNA metabarcoding chart, infographic or report that displays relative sequencing-read abundance without turning the percentage of reads directly into the percentage of animals, organisms or biomass.
It does not own molecular biology, PCR, genetics, taxonomy, ecology, population estimation or statistical modelling as standalone concepts. It uses those ideas only to show why a real scientific communication object needs a careful evidence boundary.
Why This Belongs in PSLE Science
The 2026 PSLE Science assessment framework, based on the 2023 Primary Science syllabus, includes interpreting and analysing information, evaluating observations, information and methods, considering more than one plausible explanation and communicating evidence-based reasoning.
An eDNA chart is a strong transfer object because the graph may be simple while the route from pond to bar chart contains many transformations. A learner who asks “What happened between the real world and this number?” is practising scientific inquiry at a high level without needing university-level molecular biology.
Rebuild the Evidence Object: From Pond to Bar Chart
Imagine an original composite investigation. Scientists collect three bottles of pond water. The simplified pathway is:
pond organisms → DNA enters water → water sample → DNA extraction → target DNA amplified → sequencing → reads assigned to taxa → relative-read chart
Every arrow is a transformation. Each transformation can change what the final number represents.
| Stage | Question to ask | Possible effect on final read share |
|---|---|---|
| DNA shedding | Do different organisms release similar amounts of DNA? | Some taxa may contribute more DNA per organism |
| Environmental transport | Did DNA move from where it was released? | DNA at the sample point may not match organisms exactly there |
| Sampling | Where, when and how much water was collected? | Patchy DNA can be sampled unevenly |
| Extraction | Was DNA recovered equally well? | Recovery can differ |
| Amplification | Do primers amplify taxa equally? | Some sequences can become over- or under-represented |
| Sequencing and assignment | How were reads produced, filtered and matched? | Read counts and taxonomic labels can change |
Observed, Processed, Assigned, Claimed and Inferred
| Layer | Example | What it supports |
|---|---|---|
| Environmental sample | Water collected from stated locations and times | Material available for analysis |
| DNA result | Target sequences recovered after extraction and amplification | Evidence that matching DNA entered the analytical pipeline |
| Read assignment | 60% of retained reads assigned to Species A | Species A dominated the assigned reads in that processed dataset |
| Bounded inference | Under a validated method, higher read counts may track abundance to a stated degree | Conditional ecological inference |
| Unsupported leap | 60% of animals in the pond were Species A | Not established from read share alone |
The Denominator Question: Sixty Percent of What?
Whenever you see “60%,” ask for the denominator. In our fictional chart, 60% means 60% of the retained, assigned sequencing reads in that dataset. It does not automatically mean 60% of:
- all animals in the pond;
- all fish in the pond;
- all animal biomass;
- all DNA molecules originally present in the water;
- all DNA molecules released during that week;
- all organisms living anywhere in the ecosystem.
The percentage can be mathematically correct and scientifically misread at the same time. The error occurs when the denominator silently changes.
Why DNA Shedding Breaks the Simple “One Animal = One Share” Idea
Living organisms release DNA through cells, mucus, scales, waste, reproductive material, tissue and other biological material. Different species, life stages and conditions can release DNA at different rates. A larger organism may contribute more DNA than a smaller one. A stressed or active organism may shed differently from a resting one. DNA also degrades and moves through the environment.
Therefore one animal is not automatically one equal “vote” in the final read count. Before the laboratory even starts, the biological source has already introduced variation.
Why PCR Can Change Relative Representation
Metabarcoding commonly uses primers to copy selected DNA regions before sequencing. If a primer matches one taxon’s target sequence more efficiently than another’s, repeated amplification can magnify the difference. Small process differences can therefore become large differences in read output.
This does not mean PCR is “bad.” It means the method has a known measurement job and known limitations. Scientists investigate and validate those limitations rather than pretending they do not exist.
The Important Nuance: Read Counts Are Not Always Useless
A student might over-correct and say, “Then read numbers tell us nothing about abundance.” That is also too strong.
A USGS-supported study of fish eDNA metabarcoding found statistically meaningful relationships between read numbers and aspects of population size, but the relationship was strongly influenced by dominant taxa and was weaker for non-dominant species. The study emphasised biases and spatial variation and described abundance inference as conditional rather than automatic.
So the better scientific rule is:
Do not assume read abundance equals organism abundance. Ask whether this exact method has been validated for quantitative inference, and how strong that relationship is.
Worked Case 1: The 60% Pie Slice
A report shows Species A = 60% of reads and Species B = 40%. A student states, “There were 1.5 times as many A animals as B animals.”
Evaluation: Unsupported unless the method has been validated to convert those read proportions into abundance ratios for these taxa under these conditions. The chart directly reports read share, not animal counts.
Worked Case 2: One Large Fish, Many Reads
A tank contains one large fictional fish species X and ten tiny fish of species Y. The sequencing result produces more reads for X than Y. A student concludes there must have been more X individuals.
Evaluation: The conclusion is not justified. Individual count is only one possible influence. Body size, DNA shedding, extraction and amplification can affect the read result. Direct counts in the tank would be stronger evidence about the number of individuals.
Worked Case 3: Same Pond, Different Primer Set
The same extracted DNA is analysed with two primer sets. Method 1 gives Species A 70% of reads. Method 2 gives Species A 45%. Did the pond population change between the two analyses?
Evaluation: No population change is needed to explain the difference because the same stored DNA was analysed. The changed laboratory method is a direct alternative explanation. Primer-dependent amplification can alter relative read representation.
Worked Case 4: More Reads After Rain
Species C’s read count doubles after heavy rain. A headline says, “Species C population doubled overnight.”
Evaluation: Far too strong. Rain can move water, sediment and DNA, change dilution, change where samples are collected effectively and alter environmental transport. A read-count change is an observation from the analytical pipeline; population doubling is a biological claim needing more evidence.
Worked Case 5: Three Replicate Bottles Disagree
Three water samples from nearby points give Species D read shares of 10%, 38% and 14%. A report averages them and presents 21% without showing the three values.
Evaluation: The average is mathematically valid, but it hides large spatial or technical variation. The disagreement itself is evidence. A careful reader should ask whether the variability reflects patchy DNA, sampling differences, laboratory variation or another cause before treating 21% as a stable description.
Worked Case 6: A Validated Within-Species Trend
A research team first performs controlled validation for Species E and shows that, within the tested range and method, higher eDNA measurements reliably track increases in an independently measured abundance indicator. During later monitoring, the same protocol shows a sustained rise.
Evaluation: This is stronger evidence for a bounded abundance trend because the relationship was tested for that method and species. The learner should still preserve uncertainty and avoid converting the signal into an exact organism count unless the validated model supports that conversion.
Representation Check: A Stacked Bar Can Look Like a Census
A 100%-stacked bar is visually powerful because it fills the whole width. The viewer’s brain naturally interprets it as “the whole community.” But the whole bar may actually represent the whole sequencing-read dataset after filtering.
Before reading a stacked bar as a biological composition chart, inspect the axis label, caption and method. Terms such as “relative read abundance,” “sequence reads,” “ASVs,” “OTUs” or “assigned reads” signal that the graphic is representing data produced by a molecular pipeline rather than a direct headcount.
Provenance Check: Where Did the Sequence Label Come From?
A DNA read does not arrive with a species name printed on it. Analytical software compares sequences with reference information and applies rules to assign taxonomy. If the reference database is incomplete or closely related species have difficult-to-separate sequences, some assignments can be uncertain or remain at a higher taxonomic level.
This means the provenance chain matters:
sample → extraction → amplification → sequencing → filtering → reference matching → taxonomic assignment → chart
The longer the chain, the more important it becomes to ask what each transformation contributes.
What Evidence Strengthens a Quantitative eDNA Claim?
- The sampling design covers relevant places and times.
- Replicate samples show whether the signal is stable.
- Controls check for contamination and failed detection.
- Primer performance and known taxon biases are evaluated.
- The analytical pipeline is documented.
- Read abundance is compared with an independent abundance measure.
- The relationship is validated for the taxa and range being studied.
- Uncertainty and variation are reported rather than hidden by one percentage.
- The conclusion distinguishes relative trends from exact counts.
What Weakens an Over-Broad Claim?
- Read percentage is relabelled as animal percentage without validation.
- One water bottle is used to describe an entire lake.
- Primer bias is ignored.
- Different methods are compared as though the read scales are identical.
- A change after rain or disturbance is assumed to be a population change.
- Replicates disagree strongly but only an average is shown.
- Taxonomic assignment uncertainty is hidden.
- DNA detection is treated as proof of a living organism at that exact point and moment.
Alternative Explanations: Why Did Species A’s Read Share Rise?
Suppose Species A rises from 30% to 60% of reads between two sampling dates. A larger population is one hypothesis, but alternatives include:
- Species A shed more DNA during the second period;
- water movement concentrated its DNA near the sampling point;
- other species produced less recoverable DNA;
- sample volume or location differed;
- extraction efficiency changed;
- PCR amplification behaved differently;
- sequencing depth and filtering changed;
- reference assignment differed;
- the denominator changed because other taxa’s reads changed.
This does not mean every explanation is equally likely. It means the read-share observation alone does not identify the cause. Evidence from controls, repeats, method documentation and independent biological measurements can discriminate among explanations.
How Far Can the Conclusion Travel?
A careful conclusion from the fictional chart is:
Under this sampling and metabarcoding pipeline, 60% of the retained reads in this dataset were assigned to Species A.
That result does not by itself prove that:
- 60% of individual animals were Species A;
- 60% of the biomass was Species A;
- Species A occupied 60% of the pond;
- the population grew by the same percentage as the read share;
- the organism was alive at the exact sampling point when sampled;
- all species were detected with equal efficiency.
Tempting but Invalid Reasoning
- “60% reads = 60% animals.” The denominator changed from reads to organisms.
- “More reads always means more individuals.” Shedding and laboratory bias can intervene.
- “Read counts are useless.” Some validated systems support bounded abundance inference.
- “Same sample means same result under every primer.” Method choice can alter relative amplification.
- “The bar chart shows the whole ecosystem.” It shows what the sampling and analytical pipeline captured and retained.
- “DNA detected means a live animal was exactly there.” Detection and organism location are different jobs.
PSLE-Style Transfer Case: The Pond With Known Counts
Scientists temporarily study a controlled pond where they independently know the number of fish. The fictional results are:
| Species | Known fish count | Share of assigned eDNA reads |
|---|---|---|
| A | 40% | 65% |
| B | 40% | 20% |
| C | 20% | 15% |
Question 1: Does read share equal fish-count share in this dataset? No. Species A has 65% of reads but 40% of the fish.
Question 2: Does the mismatch prove the sequencing failed? No. The method may still be excellent for detection or community comparison while requiring calibration for abundance estimation.
Question 3: What useful next investigation could be performed? Repeat samples and test whether each species has a stable relationship between read output and independently known abundance under the same method.
Question 4: Why is an independent known count powerful? It provides an external reference against which the molecular proxy can be tested.
Explained Practice
1. What does “60% of reads” directly describe? The sequencing-read dataset after the stated processing and assignment steps.
2. Name one biological reason reads may not equal individuals. Different DNA shedding rates.
3. Name one laboratory reason. Different primer amplification efficiency.
4. What should support a claim that reads estimate abundance? Validation against an independent abundance measure.
5. Why use replicate water samples? To reveal spatial or technical variation and test whether the pattern is stable.
6. Can read abundance ever be informative about abundance? Yes, in some validated contexts, but the relationship is conditional rather than automatic.
7. What is the core evidence habit? Follow the transformation chain from organism to published number.
Delayed Independent Return
Tomorrow, redraw this chain from memory:
organism → environmental DNA → sample → extraction → amplification → sequencing → taxonomic assignment → read chart
Put a question mark above every arrow where bias or uncertainty could enter. Then explain why the final bar chart is evidence, but not a direct photograph of the living community.
Useful eduKateSengkang Routes
- Reality Lab Vol No.153 | Environmental DNA Detected — Does That Prove a Living Animal Was Right There?
- Reality Lab Vol No.257 | Detected in 70% of Samples Is Not 70% of Time or Area
- Reality Lab Vol No.267 | Population Estimate Is Not a Literal Animal Count
- Reality Lab Vol No.273 | Diversity Index Is Not Simply Species Count
Parent and Tutor Teaching Guide: The Paint-Mixer Analogy—With a Warning Label
Use three coloured beads to represent three species. Place different numbers of beads into cups, but tell the learner that each bead releases a different amount of coloured dye. Then take one spoonful from the mixed liquid. The colour intensity in the spoon will not necessarily reveal the number of beads of each colour.
Now add the warning: the analogy is only a model. Real eDNA involves DNA shedding, transport, extraction, amplification and sequencing—not dye. Ask the learner to name where the analogy helps and where it stops.
This two-step exercise teaches both evidence transfer and model limits: a proxy can contain information about a hidden quantity without being identical to that quantity.
Authoritative Sources
- Singapore Examinations and Assessment Board — PSLE Science syllabus for examination from 2026
- Ministry of Education, Singapore — Science Teaching & Learning Syllabus, Primary, 2023
- U.S. Geological Survey — Environmental DNA metabarcoding read numbers and their variability predict species abundance, but weakly in non-dominant species
- U.S. Geological Survey Publications Warehouse — study record and abstract
- NOAA Institutional Repository — 12S metabarcoding with a DNA standard and quantitative limitations
The USGS study is particularly useful because it avoids a simplistic all-or-nothing story. It reports that read numbers can relate to population measures in a studied system while also showing taxon-specific biases, dominant-species effects and weaker inference for non-dominant taxa. That is exactly the scientific reasoning habit we want: preserve useful signal without pretending the proxy is the thing itself.
The Quiet Rule to Keep
A scientific pipeline can transform a real organism into a sample, a molecule, a copied sequence, a digital read and finally a coloured bar. Every transformation can add information—and every transformation can also add a boundary.
When the chart says “60% of reads,” keep the noun attached to the number. Do not silently replace reads with animals. That one habit prevents an enormous number of scientific mistakes.