Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

PSLE Science Reality Lab Vol No.108 | “The Duplicate Samples Disagreed” — Should We Average Them and Move On?

PSLE-SCI-REALITY-0108

Wait, What? The Average Can Look Calm Even When the Two Measurements Are Shouting at You

Two samples are collected from what is described as the same place at the same time. The first gives 8 units. The second gives 20 units.

Someone writes:

Average result = 14 units.

The arithmetic is correct. The scientific story may still be incomplete.

If two measurements were intended to be duplicates or replicates, their disagreement is itself evidence. It may point to natural variation in the material, differences in how the samples were collected, contamination, sample preparation, instrument variability or a mistake. Replacing 8 and 20 with one neat 14 before investigating the difference can erase the very clue the quality check was designed to reveal.

The Reality Lab habit is: when repeated evidence disagrees, examine the disagreement before summarising it.

Quick Answer

  1. Find out exactly what was duplicated: the same prepared sample, one sample split into two portions, or two independently collected samples from the same place and time.
  2. Compare the individual results before calculating an average.
  3. Ask whether the observed difference is expected for the method and sample type.
  4. Keep several explanations alive: environmental heterogeneity, sampling, transport, preparation, instrument behaviour and human error.
  5. Use averaging only after deciding what the disagreement means. A mean is a summary, not a repair tool.

The Exact Learner Job This Page Owns

This page owns one real-world evidence-transfer problem: a scientific report presents disagreeing duplicate or replicate results, and the communication hides the disagreement by reporting only their average.

It does not replace the canonical PSLE Science owners for repeated trials, variation, anomalous results, sampling or averages. It applies those skills to an actual quality-control object: a duplicate pair that was created precisely because agreement or disagreement can tell us something about the evidence chain.

Original Reality Lab Case: Two Bottles From One Stream

This is an original composite teaching case. The data are constructed for learning.

A field team collects two bottles from the same sampling point within one minute. Both are analysed for fictional Indicator P using the same method.

Field sampleIndicator P
Duplicate A8.1 units
Duplicate B19.7 units

A public infographic reports only “Mean = 13.9 units” and places a green dot next to the result.

The mean is mathematically correct. But it hides a scientific question: why did two samples intended to represent nearly the same condition differ so much?

Perhaps the stream was genuinely patchy. Perhaps one bottle collected more suspended particles. Perhaps one container was contaminated. Perhaps the sampling technique differed. Perhaps the method has large variability at this level. Perhaps a label or transcription error occurred. The difference does not tell us which explanation is correct, but it tells us not to act as though no explanation is needed.

First Question: What Does “Duplicate” Mean Here?

Scientific programmes do not always use the words duplicate and replicate in exactly the same way. That is why a strong reader asks what was physically repeated.

Repeated objectWhat disagreement can reveal
Same prepared portion measured twiceInstrument or measurement repeatability under nearly identical conditions
One sample split into two portionsPreparation and analytical variability after the split
Two samples independently collected at the same place and timeField sampling variability plus natural small-scale variation, transport, preparation and analysis

EPA quality-assurance guidance makes a similar distinction between split samples and independently collected replicate samples. The learner-level lesson is powerful: the place where the process was repeated determines what the duplicate can test.

Observed, Claimed and Inferred

LayerStatement
ObservedThe two reported values are 8.1 and 19.7 units.
CalculatedThe arithmetic mean is 13.9 units.
Scientific observationThe duplicate pair shows substantial disagreement.
Unsupported shortcut“The true value is 13.9, so the disagreement does not matter.”
Better inferenceThe disagreement is evidence that variability or an error source should be investigated before deciding how to summarise or use the result.

Why Averaging Is Not the First Scientific Move

An average is useful when repeated measurements are reasonably understood as samples from a common process and when the spread among them is compatible with the purpose of the measurement. But an average does not explain disagreement.

Consider two pairs:

PairResultsAverage
X13.8 and 14.013.9
Y8.1 and 19.713.9

The averages are identical. The evidence patterns are not. Pair X suggests close agreement. Pair Y tells us that the single summary number hides a large spread.

This is why a reader should inspect the individual evidence before accepting a summary statistic.

The Representation Check: Did the Graphic Show the Pair or Only the Mean?

Communication design can make disagreement disappear. A bar chart with one bar at 13.9 units looks calm. Two points at 8.1 and 19.7 tell a different story.

  1. Are individual duplicate results visible?
  2. Is their difference shown numerically or graphically?
  3. Does the chart display an average without the underlying spread?
  4. Were any duplicates excluded before the summary was calculated?
  5. Does the caption explain whether the disagreement met the project’s quality criterion?

The graphic should help readers see the evidence structure rather than smooth it away.

The Comparison Check: Is the Difference Large Enough to Matter?

Not every difference between duplicates is alarming. No real measurement system produces perfectly identical values forever. The key question is whether the disagreement is acceptable for the method, sample type and scientific purpose.

A Primary 5/6 learner does not need to memorise laboratory formulas or universal percentage limits. In fact, there is no single tolerance that belongs to all measurements. Strong reasoning asks whether the project has a justified criterion and what happens when a duplicate pair falls outside it.

Alternative Explanation 1: The Environment Was Actually Patchy

Two field samples can disagree because the system itself is not perfectly mixed. Sediment may collect in patches. Tiny organisms may cluster. Suspended particles may move unevenly. Water near a surface can differ from water a little deeper.

If natural heterogeneity is the cause, the disagreement is not merely “measurement error”. It is evidence about the system. More sampling may be needed to describe that variability honestly.

Alternative Explanation 2: The Sampling Procedure Added Variability

One bottle may have been filled differently, collected at a different depth, mixed more strongly, exposed to a dirty cap or held for a different length of time. Field duplicates can reveal this combined variability because they repeat more of the real sampling process than a repeated instrument reading does.

Alternative Explanation 3: Preparation or Analysis Varied

If one original sample is split into two portions and the portions later disagree, the source of variation is likely to be downstream of the split—or in imperfect mixing before the split. That narrows the investigation.

Alternative Explanation 4: One Result Is Wrong

A label swap, transcription error, contaminated container, instrument fault or calculation mistake can happen. But disagreement alone does not tell us which result is wrong. Deleting the inconvenient value without evidence is no better than averaging it away.

The correct move is diagnosis: inspect records, blanks, calibration checks, sample history and repeat evidence.

What Evidence Would Strengthen the Result?

  • A clear definition of what the duplicate or replicate actually repeats.
  • A stated, method-appropriate criterion for acceptable agreement.
  • Additional replicate samples collected independently when natural variability is suspected.
  • Clean blanks and appropriate reference checks that reduce some alternative explanations.
  • Documented sample handling showing both members of the pair followed comparable conditions.
  • A repeat analysis or independent method when a possible analytical problem needs testing.
  • Raw individual values reported alongside any average.

What Would Weaken It?

  • Only the average is shown even though duplicate disagreement was large.
  • The report never explains what was duplicated.
  • One member of the pair is deleted because it “looks wrong” with no diagnostic evidence.
  • Different storage or collection conditions are ignored.
  • A single duplicate pair is used to claim that the entire environment is uniform.
  • The project has no plan for what to do when duplicates disagree.

Worked Case 1: Same Aliquot, Two Instrument Readings

The same prepared liquid is measured twice within a minute. Results are 12.1 and 12.2 units. The close agreement supports good short-term repeatability under those measurement conditions. It does not prove that the sample represents the whole pond or that 12.15 is the exact true value.

Worked Case 2: One Bottle Split in Two

A well-mixed bottle is divided into Portion A and Portion B before separate preparation. Results are 11 and 18 units. Because both portions came from the same bottle, large disagreement points attention toward mixing, preparation or analytical variability after the split. It tells us less about whether the original pond was patchy.

Worked Case 3: Two Independent Field Samples

Two bottles are independently collected side by side. Results are 7 and 15 units. This pair includes more sources of variation: natural small-scale differences, sampling technique, transport, preparation and analysis. The correct next step is to identify which source matters rather than call one bottle “wrong” automatically.

Worked Case 4: The Average That Changes the Headline

A project threshold for a broad communication claim is 14 units. Duplicate results are 8 and 20, so the average is 14. Reporting only “14” makes the evidence appear to sit exactly on the boundary. Reporting both values reveals that the pair straddles it widely. That difference may affect how cautious the conclusion should be.

Tempting Reasoning That Fails

  • “The average is always more scientific than the individual results.” An average can hide important spread.
  • “One of two different results must be wrong.” Natural variability can produce genuine differences.
  • “Duplicates disagree, so the experiment failed.” Disagreement may be exactly the information the duplicate was designed to reveal.
  • “Take a third measurement and keep whichever two agree.” That can introduce selection bias unless there is a justified rule.
  • “Close duplicates prove accuracy.” Two measurements can agree closely and still share the same systematic bias.

Model and Measurement Limits

The meaning of duplicate disagreement depends on how the duplicates were created. Laboratory terminology also varies among programmes. That is why this page deliberately avoids a universal numerical cutoff. The scientific skill is to reconstruct the repeated process, then ask whether the difference is reasonable for the method and purpose.

A duplicate pair also provides limited information about the full distribution of variability. Two points can reveal a problem but cannot fully describe every source of variation in a complex system.

How Far Can the Conclusion Travel?

Close duplicates can strengthen confidence that the repeated portion of the measurement process is reasonably consistent. Widely separated duplicates tell us that variability deserves attention. Neither result, by itself, establishes representativeness across a large environment, proves causation or makes a broad product claim universal.

PSLE-Style Transfer Case

A pupil measures the mass of two supposedly identical portions prepared from the same mixture. The values are 24 g and 39 g. Another pupil says, “The average is 31.5 g, so just use that.”

Question: What should be checked before accepting 31.5 g as the useful result?

Reasoned answer: The large disagreement should be investigated first. The portions may not have been prepared equally, the mixture may not have been well mixed, or a measurement error may have occurred. Averaging does not explain why the two values differ.

Explained Practice

Practice A: Duplicate readings are 10.0 and 10.1 units. What does the pair mainly tell you? The repeated measurement has close agreement under those conditions. It does not alone establish accuracy.

Practice B: Field replicates are 4 and 17 units. What should happen before reporting their mean as a site value? Check the sampling method, local variability, sample histories, QC records and whether more representative sampling is needed.

Practice C: A report displays one mean but notes “duplicate failed acceptance criterion”. Should the warning matter? Yes. The summary should not erase a quality-control result that changes how confidently the number can be used.

Delayed Independent Return: The T-W-I-N Check

  1. T — Twin of what? What part of the process was actually repeated?
  2. W — Width of disagreement: How far apart are the individual results?
  3. I — Investigate: Which plausible sources of variation or error could explain the difference?
  4. N — Next evidence: What additional sample, control, record or repeat would distinguish those explanations?

Parent and Tutor Teaching Guide

Write two result pairs on cards: 13.8 and 14.0; 8.1 and 19.7. Ask the learner to calculate both averages. When the learner discovers that both averages are about 13.9, ask: “Do these two evidence stories deserve the same confidence?” That single contrast makes the purpose of variability visible.

Next, change what was duplicated. First repeat the same instrument reading. Then split one sample. Then imagine collecting two new samples side by side. Ask which sources of variation each design includes. This helps the learner see experimental design as a map of possible explanations rather than a vocabulary exercise.

Authoritative Sources

EPA guidance distinguishes several kinds of QC samples because they reveal different parts of the evidence chain. USGS quality-assurance programmes similarly use replicates, blanks and spikes to monitor field and laboratory performance. The official Singapore Science frame then supplies the learner habit: interpret information, evaluate observations and methods, consider uncertainty and communicate reasoning.

The Quiet Return

Two disagreeing measurements are not an inconvenience to be tidied away.

They are a question the experiment has asked you.

Before averaging repeated evidence, find out what the disagreement is trying to teach you.