Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

PSLE Science Reality Lab Vol No.472 | “30× Genome Coverage” — Was Every DNA Base Read Exactly 30 Times?

Thirty times. Every letter. No exceptions? A genome-sequencing report says 30× coverage. A student imagines a careful machine passing over every DNA position exactly thirty times, like a teacher checking every word in an essay with thirty identical rereads.

That picture is too exact. Sequencing coverage is evidence about how much overlapping sequence information was obtained, but coverage depth can vary from one genomic position to another. A report describing a genome as “30×” commonly summarises coverage across a much larger region. It does not, by those two characters alone, guarantee that every single base was read exactly thirty times, that no position had low coverage, or that every read was equally reliable.

Quick answer

No. “30× genome coverage” should not be read as “every DNA base was observed exactly 30 times.” The correct learner job is to ask what the 30× number summarises, how coverage is distributed across positions, whether important regions have enough usable reads, what quality checks were applied, and whether the scientific claim depends on parts of the genome that were weakly covered.

The owned learner job — and what belongs elsewhere

This Reality Lab applies PSLE Science evidence reasoning to one communication object: a sequencing report, methods statement or infographic that displays a coverage value such as 30×. It does not teach genetics, inheritance, DNA replication, variant calling, sequencing chemistry or bioinformatics as standalone concepts. It also does not turn Primary learners into clinical interpreters.

The evidence habit belongs with the existing owners. Use Observation, Inference, Prediction and Explanation to separate measured evidence from a conclusion, and How to Evaluate PSLE Science Observations, Information and Methods for the general method-checking skill. This page owns only the transfer question: what does a coverage summary allow us to say about the underlying sequencing evidence?

Build the idea with a tiny invented genome

Imagine a completely fictional stretch of DNA with only eight positions: A, B, C, D, E, F, G and H. After sequencing and alignment, the number of usable reads covering each position is:

PositionReads covering it
A28
B31
C35
D29
E42
F17
G33
H25

The mean depth of this small example is 30 reads per position. Yet none of the eight positions has to be exactly 30. One position has 42; another has only 17. A single summary number can be mathematically correct while hiding unevenness that matters for a particular claim.

Coverage is a map of support, not a stamp

The useful mental image is not “thirty perfect passes.” Think instead of many overlapping strips laid across a long reference line. Some places receive more overlaps, some fewer. The depth at one position asks: how many aligned reads support information here?

NHGRI explains depth of coverage using repeated sequencing of genomic regions and explicitly describes values such as 1× or 4× in average terms. More recent NHGRI teaching material shows coverage rising and falling along a contig and notes that sequencing reads are sampled in ways that require oversampling to build dependable coverage. The important learner move is to read “30×” as a summary with a distribution underneath it.

Observed, processed, summarised and inferred

LayerExampleOverreach to avoid
Observed/generatedSequencing instruments produce many reads with quality informationAssuming every read is equally reliable
ProcessedReads are filtered and aligned or assembledAssuming every generated read contributes to every base
SummarisedA report states “30× average depth”Assuming every base has depth 30
EvaluatedCoverage distribution and low-depth regions are checkedAssuming the average alone describes all local evidence
InferredA sequence or variant claim is madeAssuming the same evidence strength applies everywhere

Two different ideas can hide behind the word “coverage”

Scientific communication can use “coverage” in related ways. Depth concerns how many reads overlap a position. Another question is breadth: what fraction of the target region has any coverage, or meets a stated minimum depth. A report can therefore have a high average depth while still containing regions with weak or missing support.

You do not need advanced genetics to evaluate the claim. You need a familiar PSLE Science habit: identify exactly what was measured, exactly what the summary describes, and exactly which conclusion is being asked to travel beyond it.

Worked case 1: same average, different evidence pattern

Two fictional sequencing runs both report an average depth of 30×.

RunLocal depths at six positionsAverage
A28, 31, 30, 29, 32, 3030×
B5, 8, 12, 45, 50, 6030×

For a claim about a position with depth 5, Run B does not gain thirty supporting reads merely because the genome-wide average is 30×. The two datasets share an average but not a distribution.

Worked case 2: one very deep region can pull up the mean

A sequencing summary contains many positions around 20× but one region around 200×. The overall mean can rise. A headline then says, “The genome was deeply covered everywhere.”

The mean alone cannot justify “everywhere.” To support that wording, the report would need evidence about the coverage distribution, such as the proportion of relevant bases above a chosen depth or a plot showing local coverage. This is the same reasoning you use whenever an average is asked to stand in for every member of a group.

Worked case 3: the important position is the weak one

A science-news graphic says a sample has “30× whole-genome sequencing.” The biological claim depends on one small location, but that position happens to have only 6 usable reads after filtering. The genome-wide label is not useless; it simply is not the most relevant evidence for that local claim.

Ask for the evidence at the location that matters. Relevant evidence beats impressive global numbers.

Worked case 4: generated reads are not the same as usable reads

Suppose an instrument generates many short reads. Some fail quality filters; some cannot be confidently placed; some are duplicates or otherwise handled by the analysis pipeline. A promotional sentence that turns “lots of reads generated” directly into “every base is strongly confirmed” skips processing steps that determine what evidence remains usable.

The learner need not memorise every pipeline. The transferable question is simple: which observations actually contributed to this result after the stated method was applied?

Worked case 5: 30× is not 30 independent experiments in every sense

Thirty overlapping reads at one position can provide repeated evidence, but it is unsafe to leap from “30 reads” to “30 fully independent confirmations.” Reads may share library preparation, sample handling, instrument conditions, alignment assumptions and other common sources of error. Repetition can reduce some random errors without eliminating every possible systematic error.

That distinction mirrors many PSLE Science investigations: repeating a measurement can strengthen evidence, yet repeating the same method under the same hidden bias does not automatically remove the bias.

Representation check: a single number versus a coverage plot

A report can communicate coverage at several levels:

  • One-number summary: fast to read, but compresses variation.
  • Histogram or distribution: shows how common different depths are.
  • Coverage-by-position plot: reveals local high and low regions.
  • Threshold summary: may report the fraction of bases at or above a stated depth.

When the claim is local, local evidence matters. When the claim concerns the whole target, distributional evidence matters. A good representation should match the scale of the conclusion.

The denominator check

If a report states, “95% of target bases had at least 20× coverage,” ask what counts as the target. Is it a whole genome, a selected panel, an exome, a particular set of regions, or only positions that passed earlier filters? The percentage is interpretable only when its denominator is known.

This is not a genetics trick. It is the same evidence habit used for percentages in environmental samples, product tests and survey results: percent of what?

What strengthens a sequencing claim?

  • The report defines whether 30× is mean, median, target or minimum coverage.
  • Coverage distribution is shown rather than hidden behind one summary.
  • Important regions have adequate local usable depth for the specific claim.
  • Read quality and mapping or assembly quality are checked.
  • Low-coverage or uncertain regions are identified instead of silently treated as equally certain.
  • The conclusion is confirmed by an appropriate independent method when the scientific question requires it.

What weakens the claim?

  • “30×” is displayed without saying what statistic or target it refers to.
  • The claim concerns a weakly covered position but cites only the genome-wide average.
  • Large low-depth gaps are hidden by a high-depth region that raises the mean.
  • Low-quality reads are counted as if they were equally trustworthy evidence.
  • The article turns sequencing depth into a universal accuracy percentage.
  • Repeated reads are treated as if they remove all possible systematic errors.

How far can the conclusion travel?

A clearly defined 30× average can support a statement about the overall amount of sequencing depth under that method. It cannot, by itself, support “every base has exactly 30 observations,” “every position is equally certain,” or “the sequence is guaranteed correct.” The narrower the scientific claim becomes, the more important local coverage and other quality evidence become.

Tempting reasoning that does not survive

  • “30× means exactly thirty reads everywhere.” Coverage depth can vary by position.
  • “30× means 30%.” The × symbol here indicates fold/depth, not percentage.
  • “The average is high, so there cannot be gaps.” A mean does not prove a uniform distribution.
  • “Thirty reads means zero chance of error.” More repeated evidence can help with some errors but does not eliminate every source of uncertainty.
  • “Two reports both say 30×, so their evidence quality is identical.” Distribution, read quality, method and target can differ.

Original PSLE-style transfer case: the tile scanner

This is an original transfer exercise, not an examination-board question.

A museum uses a scanner that photographs overlapping strips of a long painted wall. The report says the wall has an average image coverage of 10×. Most tiles appear in 9–12 photographs, but one narrow area appears in only 2 usable photographs and another in 28. A student says, “Every tile has ten photographs because the report says 10×.”

Question 1: Why is the student’s statement incorrect?

Explained answer: The 10× value is an average summary. Individual positions can have different numbers of overlapping observations.

Question 2: If a claim concerns a tiny crack in the 2× area, which evidence matters most?

Explained answer: The local usable images at the crack matter more than the average coverage of the whole wall.

Question 3: Would increasing the average to 20× guarantee every tile has at least 20 photographs?

Explained answer: No. A larger average still does not prove a minimum at every location unless the distribution or minimum is separately reported.

Delayed independent return

  1. What is the difference between an average coverage depth and the depth at one DNA position?
  2. Why can a high average hide low-coverage regions?
  3. Why does a claim about one location require local evidence?
  4. Why are generated reads not automatically the same as usable supporting reads?
  5. Why does repetition reduce some uncertainty without guaranteeing perfection?

Self-check: summary and local value are different; distributions can be uneven; relevant evidence must match the claim; quality and processing determine what counts as usable; repeated observations help but do not erase systematic limitations.

Practice: read the fine print under “30×”

  1. A report says “mean depth 30×; 92% of target bases ≥20×.” Which statement is safer: “every base has 30 reads” or “most target bases met the stated 20× threshold”? The second.
  2. Two bases have depths 3 and 70 in a 30× dataset. Can the average tell you which is better supported? No; inspect the local depths and quality.
  3. A chart shows a large region with no aligned reads. Can a 30× genome-wide mean fill that evidence gap? No.
  4. A headline says “30× coverage means 30 times more accurate.” Is that justified? No; depth is not a universal linear accuracy multiplier.

Parent and tutor teaching guide: use transparent strips

Cut ten transparent strips of different lengths and lay them over a printed line divided into twenty boxes. Ask the learner to count how many strips cover each box. Then calculate the average. The learner will immediately see that the mean can be, say, 5× even when some boxes are covered twice and others eight times.

Next ask three questions in order: “What does the average tell us?”, “What does it hide?”, and “Which box matters for the claim?” That sequence keeps the lesson about evidence, not genetics vocabulary.

For three students, assign roles: summary reader states only what the headline statistic says; distribution reader finds unevenness; claim checker decides what additional local evidence is needed. Rotate roles after each case.

Authoritative sources and curriculum frame

The quiet habit to keep

When a scientific report compresses millions or billions of observations into one impressive number, do not throw the number away. Expand it. Ask what was averaged, how uneven the underlying evidence is, and whether the part that matters for the claim is well supported. “30×” becomes informative when you know what lies underneath the multiplication sign.