Reality Lab ID: PSLE-SCI-REALITY-0499
Wait, what? A sequencing dashboard shows a large green badge: Q30. One learner says, “That means the DNA was read thirty times.” Another says, “No, Q30 must mean thirty percent accurate.” A third remembers seeing “30× coverage” in another article and assumes the two thirties are probably the same. They are not.
This PSLE Science Reality Lab is about a modern scientific communication object: the quality score printed beside DNA sequencing data. Primary 5 and Primary 6 learners do not need to learn molecular diagnostics or become genome analysts. They do need to practise a durable evidence habit: when a number appears in a scientific report, identify the quantity, scale, denominator and method before attaching an everyday meaning to it.
The current 2026 PSLE Science framework assesses not only scientific knowledge but also the application of knowledge and scientific inquiry, including interpreting and analysing information, evaluating observations, information and methods, and communicating reasoning. The 2023 Primary Science syllabus also values integrity, objectivity, open-mindedness and healthy scepticism. A Q score is useful practice because its number looks simple while its meaning comes from a mathematical scale and a specific stage of the measurement process.
Quick Answer
Q30 does not mean 30× coverage, and it does not mean 30% accuracy. A Phred-like quality score expresses an estimated probability that a particular base call is wrong on a logarithmic scale. Under the standard definition, Q30 corresponds to an estimated error probability of 1 in 1,000 for that base call, which is equivalent to an inferred base-call accuracy of 99.9% under the scoring model.
But even that sentence needs scope. A Q30 score for one base does not guarantee the entire DNA read, sample or genome is error-free. A dashboard that says “90% bases ≥ Q30” is reporting the fraction of base calls meeting a threshold. A report that says “30× coverage” is describing how much sequencing overlaps positions in the target. Those are different quantities answering different questions.
The Owned Learner Job
This article owns one job: how to evaluate a sequencing quality label such as Q30 without confusing a logarithmic base-call error score with coverage, percent accuracy for the whole experiment, or proof that a biological conclusion is correct.
It does not re-teach DNA structure, genetic inheritance, probability or logarithms as standalone science owners. It also does not provide medical interpretation. For the wider evidence system, route to How Scientific Evidence Works, How to Keep a PSLE Science Claim at the Right Evidence Level, and How to Tell Precision From Accuracy.
First Decode the Letter
The letter Q is not a unit like metres, grams or seconds. In this context it refers to a quality score derived from an estimated probability of a base-calling error. A sequencing machine and its software observe signals generated during sequencing and decide which DNA base—A, C, G or T—best fits the evidence. The software can also estimate how uncertain that call is.
That uncertainty is converted into a score. The standard Phred form is:
Q = −10 log10(p), where p is the estimated probability that the base call is wrong.
You do not need to calculate logarithms for PSLE Science. What matters is understanding the scale: every increase of 10 Q units corresponds to a tenfold decrease in the stated error probability.
| Quality score | Estimated error probability | Equivalent inferred base-call accuracy |
|---|---|---|
| Q10 | 1 in 10 | 90% |
| Q20 | 1 in 100 | 99% |
| Q30 | 1 in 1,000 | 99.9% |
| Q40 | 1 in 10,000 | 99.99% |
Notice that Q20 to Q30 is not a change from 20% to 30%. It is a tenfold change in the estimated error probability. The score is logarithmic, so reading the numeral as a percentage breaks the scale.
Original Composite Case: The River-Reed Genome Run
Imagine a fictional school research project sequencing DNA from a harmless river reed. The dashboard reports:
| Dashboard field | Value |
|---|---|
| Total bases read | 12,000,000 |
| Bases at or above Q30 | 10,800,000 |
| Percent bases ≥ Q30 | 90% |
| Average target coverage | 30× |
| Median read length | 150 bases |
There are two “30” labels: Q30 and 30×. They are unrelated in meaning.
- Q30 is a base-call quality threshold tied to estimated error probability.
- 30× coverage describes how much sequencing data, on average, overlap the positions in the target sequence.
The dashboard also says 90% bases ≥ Q30. That is a third quantity. It tells us that 90% of the base calls reached at least the Q30 threshold. It does not mean the sample has “Q90”, and it does not mean every base is Q30.
Observed, Called, Scored, Summarised
A sequencing result passes through several layers. Keep them separate:
- Observed signal: the instrument detects physical or chemical signals generated during sequencing.
- Base call: software interprets a signal as A, C, G or T.
- Quality score: software estimates the probability that a particular call is wrong and converts it to a Q score.
- Run summary: the report may count what fraction of millions of bases meet a threshold such as Q30.
- Biological interpretation: scientists later use the sequence data to answer a biological question.
A high-quality base call supports the first steps of that chain. It does not automatically prove every later interpretation. That is a recurring scientific lesson: evidence quality at one stage strengthens later work but does not erase the need to check later stages.
Q30 Is Usually About a Base Call, Not a Whole Genome
The classic Phred idea assigns quality to individual called bases. If one base is Q30, that score describes the estimated error probability for that call under the scoring system. A full read contains many base calls. A sequencing run contains millions or billions of them.
Therefore a report may summarise an enormous set of scores. “90% of bases ≥ Q30” means ninety percent of the base calls passed that threshold. It is not scientifically correct to say, “The whole genome is Q30, therefore there can be no mistakes anywhere.”
The Difference Between Q30 and 30× Coverage
Coverage asks how many sequencing observations overlap a target position, usually summarised across many positions. Quality asks how reliable a particular base call is estimated to be.
Imagine reading a sentence copied by many students. Coverage is like asking how many independent copies include a particular word. Base quality is more like asking how confident you are that one student copied one letter correctly. More copies can help cross-check uncertain letters, but the two ideas are not the same measurement.
The Reality Lab already has a separate owner for this other quantity: Vol No.472 on “30× genome coverage”. That article owns coverage distribution. This article owns Q-score interpretation.
Why Q30 Does Not Mean “One Error Every 1,000 Bases” as a Guarantee
Q30 corresponds to an estimated probability of 0.001 that a particular base call is wrong under the scoring model. It is tempting to turn that into a clockwork promise: exactly one wrong base in every block of 1,000.
Probability does not work that way. Across many suitable base calls, the expected error frequency may be around the stated rate if the scores are well calibrated. But one particular block of 1,000 bases could contain zero errors, one error or several. The score expresses uncertainty, not a schedule.
Calibration Check: Does the Score Match Real Error Rates?
A useful quality score should be calibrated: calls assigned a certain error probability should, over many known examples, show error rates reasonably consistent with that probability. Scientific papers can compare predicted quality with observed errors using reference data.
This matters because a score is produced by an algorithm. The formula gives the meaning of the number, but the accuracy of the estimated probability depends on how well the model fits the real data. A Q30 label should not be treated as magic independent of the platform, software and calibration evidence.
Distribution Check: One Average Can Hide Weak Regions
Suppose Read A has quality scores near Q35 across nearly all its bases. Read B begins at Q38 but drops steadily to Q15 near the end. Both could have a similar average under some summary, yet the pattern is different.
A run-level percentage such as “92% ≥ Q30” can also hide where the lower-quality bases occur. Are they clustered at the ends of reads? Concentrated in one cycle? Associated with a particular sequence context? A summary is useful, but the shape of the data may matter for a scientific decision.
Threshold Check: What Does “≥ Q30” Mean?
The symbol “≥” means “greater than or equal to”. If a dashboard says 90% of bases ≥ Q30, the included bases could have Q30, Q31, Q35, Q40 or higher scores. The label does not say all 90% are exactly Q30.
The remaining 10% are not automatically “wrong”. They simply fall below the chosen Q30 threshold. A Q25 call still has a quality score; it is not equivalent to no information. Whether it is acceptable depends on the analysis and quality-control rules.
Q30 Does Not Prove the Sample Identity
High base-call quality means the instrument/software is estimated to have called bases reliably. It does not by itself prove the correct sample was sequenced. A perfectly read wrong sample is still the wrong sample.
Sample identity depends on labelling, chain of custody, preparation, controls and other evidence. This is a useful transfer lesson: measurement quality and object identity are separate claim types.
Q30 Does Not Prove a Biological Conclusion
Suppose a high-quality DNA sequence is later compared with a database. The final species identification can still depend on reference quality, similarity thresholds, contamination checks, alignment and biological interpretation. Q30 strengthens confidence in a base call. It does not directly certify every downstream conclusion.
The Reality Lab has a separate owner for one such downstream communication object: Vol No.324 on a 99% DNA barcode match. Do not merge a base-call quality score with a species-match percentage.
Evidence That Strengthens a Sequencing-Quality Claim
- The report defines exactly what Q score is used.
- It states whether the number is per base, an average, a median or a percentage above a threshold.
- Quality calibration is checked against known reference material or control data.
- The distribution of quality across cycles or read positions is available when relevant.
- Coverage is reported separately rather than mixed into the quality score.
- Sample and process controls show that the data came from the intended workflow.
- The software version and processing steps are documented.
Evidence That Weakens an Overconfident Claim
- A dashboard shows “Q30” with no explanation of whether it means one base, a threshold percentage or an average.
- A writer confuses Q30 with 30× coverage.
- The report converts Q30 into “100% correct”.
- Quality scores are used to prove sample identity or biological meaning without independent checks.
- Only a run average is shown despite a strong drop in quality across some read positions.
- The platform’s predicted scores have not been checked against observed error behaviour.
How Far Can the Q30 Conclusion Travel?
If one base has Q30, you may say the scoring model assigns it an estimated error probability of about 0.1%, equivalent to 99.9% inferred base-call accuracy. If a run has 90% of bases ≥ Q30, you may say 90% of its base calls meet or exceed that quality threshold.
You may not jump directly to “90% of the genome is correct”, “the specimen is definitely Species A”, “the experiment has 30× coverage”, or “the biological conclusion has a 0.1% chance of being wrong”. Those are different claims with different evidence requirements.
Worked Case 1: Q20 vs Q30
A student says Q30 is only 50% better than Q20 because 30 is 50% larger than 20.
Repair: The Q scale is logarithmic. Q20 corresponds to an estimated error probability of 1 in 100; Q30 corresponds to 1 in 1,000. The error probability is ten times smaller, not merely 50% smaller.
Worked Case 2: “95% Q30”
A run report says “Q30 = 95%”. A student says each base has a 95% chance of being correct.
Repair: In many dashboards, that wording means 95% of base calls have quality scores at or above Q30. It does not mean each base has 95% accuracy. Read the metric definition.
Worked Case 3: Q30 and 30× in One Table
A sample has 88% bases ≥ Q30 and average coverage 30×. Can we say “the sample scored 30 twice”?
Repair: No. One “30” belongs to a logarithmic base-quality threshold; the other belongs to the average number of overlapping sequencing observations. Identical numerals do not imply identical quantities.
Worked Case 4: High Q Score, Wrong Tube
A sequencing run has excellent Q scores, but a sample-label check later shows the wrong tube was loaded. Does high Q rescue the result?
No. High Q supports the reliability of base calls from the material that was sequenced. It does not establish that the intended sample was in the tube. Provenance and measurement quality are separate.
Worked Case 5: One Q30 Base in a Weak Read
A 150-base read contains one Q30 base and many Q12 bases. A headline says “This read is Q30 quality.” Is that supported?
Not from that evidence. One Q30 base does not define the whole read. A read-level summary would need a stated rule, such as average score, minimum score or percentage above threshold.
Worked Case 6: 1 in 1,000 Is Not a Schedule
A student expects exactly one error in every consecutive block of 1,000 Q30 base calls.
Repair: The score expresses estimated probability. Over large sets, observed error frequency may approach the expected rate if the score is calibrated, but errors do not have to appear at perfectly regular intervals.
Tempting but Invalid Reasoning
- “Q30 = 30% accurate.” The scale is logarithmic; Q30 corresponds to about 99.9% inferred base-call accuracy.
- “Q30 = thirty reads.” That confuses quality score with coverage.
- “Q30 means no errors.” It represents a small estimated error probability, not zero.
- “90% bases ≥ Q30 means the genome is 90% correct.” It means 90% of base calls meet the stated quality threshold.
- “High Q proves the sample identity.” Quality and provenance are separate.
- “High Q proves the scientific conclusion.” Downstream analysis has its own evidence chain.
PSLE-Style Transfer Case
A fictional sequencing report shows: “92% bases ≥ Q30; average coverage = 18×.” A student writes, “The DNA was read 30 times and 92% of the genome is guaranteed correct.” Evaluate the statement.
Strong answer: The statement mixes two different metrics. Q30 is a base-call quality threshold corresponding to an estimated error probability of about 1 in 1,000 for a Q30 call; it does not mean the DNA was read 30 times. The report separately says average coverage is 18×. Also, “92% bases ≥ Q30” means 92% of base calls meet or exceed the Q30 threshold, not that exactly 92% of the whole genome is guaranteed correct. Other errors and analysis steps must still be considered.
Practice 1: Convert the Meaning, Not the Number
Q20 corresponds to what estimated error probability?
Answer: About 1 in 100, or 1%, under the standard Phred definition.
Practice 2: Ten Q Units Higher
What happens to the stated error probability when a quality score rises from Q20 to Q30?
Answer: It becomes ten times smaller: from about 1 in 100 to 1 in 1,000.
Practice 3: Threshold Percentage
A dashboard says 85% bases ≥ Q30. Does that mean the other 15% are all wrong?
Answer: No. They are below the Q30 threshold, but many can still have useful quality scores and can still be correct.
Practice 4: Q30 vs Coverage
Which metric would you inspect to ask how many reads overlap a DNA position?
Answer: Coverage or depth, not the Q score.
Practice 5: One Base vs Whole Read
Can the Q score of one base be copied onto an entire 150-base read?
Answer: Not without a stated read-level summary rule. Individual bases can have different quality scores.
Practice 6: Quality vs Identity
Could a wrongly labelled sample still produce very high Q scores?
Answer: Yes. The instrument can read the wrong sample very accurately. Sample identity needs separate evidence.
Practice 7: Quality vs Interpretation
If every base in a short DNA fragment has high Q scores, does that alone prove the organism’s species?
Answer: No. Species identification depends on what sequence was obtained, reference comparisons, method validity and other evidence.
Practice 8: The Right Claim Scope
Write a cautious sentence for “93% bases ≥ Q30”.
Answer: “Ninety-three percent of the reported base calls met or exceeded the Q30 quality threshold under this sequencing system.”
Delayed Independent Return
Without looking back, create four cards: Q20, Q30, 20× coverage and 30× coverage. Sort them into quality score and coverage. Under Q20 and Q30, write the error probabilities. Under 20× and 30×, write what the “×” refers to.
On another day, invent a dashboard with “95% bases ≥ Q30” and “25× coverage”. Explain it aloud without using the word “accurate” until you have named exactly what object the accuracy statement refers to. That forces the scope to stay visible.
Parent and Tutor Teaching Guide
Start with two bags of letter tiles rather than DNA. Tell the learner that one copying method has an estimated one-letter error in 100, while another has one in 1,000. Ask which is more reliable and whether the second method guarantees zero errors. Then introduce the labels Q20 and Q30 only after the probability meaning is secure.
Next, create a separate stack of twenty copies of one sentence. Call that “20× coverage” for one word position. This physical separation helps the learner see that repeated observations and confidence in each observation are different axes.
Finally, give a deliberately confusing card: “90% ≥ Q30; 30× coverage”. Ask the child to underline the percent sign, Q and × symbol in different ways and state the denominator or reference object for each. Scientific literacy often begins by refusing to let similar numerals erase different units.
Routes to Existing PSLE Science Owners
- How Scientific Evidence Works — for observation-to-claim structure.
- How to Tell Precision From Accuracy — for the general measurement distinction.
- How to Keep a Claim at the Right Evidence Level — for not stretching one metric into a larger conclusion.
- Reality Lab Vol No.472 — the separate owner for 30× genome coverage.
- Reality Lab Vol No.324 — the separate owner for DNA barcode match percentages.
Authoritative Sources and Further Reading
- Singapore Examinations and Assessment Board — 2026 PSLE Science syllabus.
- Ministry of Education, Singapore — 2023 Primary Science syllabus.
- Illumina — Sequencing Quality Scores, including the standard Q-score relationship between error probability and score.
- U.S. National Library of Medicine / PMC — discussion of Phred quality scores in base calling.
- U.S. National Library of Medicine / PMC — estimating Phred scores and base-call error probabilities.
The Quiet Return
A scientific score is not the numeral printed on the screen. It is the definition behind the numeral.
When you see Q30, ask what is being scored, what probability the scale represents and whether the report is talking about one base, many bases or a run-level threshold. Then keep coverage and biological meaning in their own lanes. That is how a small label becomes a serious lesson in scientific evidence.