Series ID: PSLE-SCI-REALITY-0443
Wait, What? Three Laboratories Can Agree and Still Be Wrong Together
Imagine that three laboratories test the same carefully prepared liquid. Laboratory A reports 52.0 mg/L. Laboratory B reports 52.1 mg/L. Laboratory C reports 51.9 mg/L. The results are astonishingly close. A headline appears: “Three laboratories agree, proving the true concentration is about 52 mg/L.”
That sounds persuasive because independent agreement is often useful scientific evidence. But now add one fact: the liquid is a reference material whose best-supported reference value is 50.0 mg/L. The laboratories agree with one another very well, yet as a group they sit about 2 mg/L above the reference. Agreement has told us something important about reproducibility. It has not, by itself, proved accuracy.
This distinction is a powerful PSLE Science habit. The current 2026 PSLE Science assessment objectives include interpreting and analysing information, evaluating observations, information and methods, and communicating explanations and reasoning. The 2023 Primary Science syllabus also encourages healthy scepticism: question the evidence carefully without automatically rejecting it. In a real-world laboratory comparison, that means asking two separate questions: Do the measurements agree with one another? and Do they agree with an appropriate reference for what is actually being measured?
Quick Answer
No. Agreement among laboratories is not automatically proof that the shared result is accurate. Close results can show good reproducibility under the comparison conditions. Accuracy needs another line of evidence: for example, a suitable reference value, a well-characterised reference material, a trustworthy comparison method, or another justified way to connect the measurement to the quantity being claimed.
- Agreement asks: How close are the results to one another?
- Accuracy asks: How close is the result to an appropriate reference or accepted value for the quantity?
- Shared bias is possible: laboratories can use the same mistaken assumption, calibration, conversion or method and therefore agree in the same wrong direction.
- Disagreement also needs diagnosis: one different result is not automatically the bad one.
- Comparison design matters: the sample must remain suitable, the laboratories must know what they are comparing, and the data must be read on the correct basis.
The compact habit is: agreement is evidence about agreement; it becomes evidence about accuracy only when the comparison is tied to a suitable reference and a sound measurement chain.
The Exact Learner Job — and What This Guide Does Not Own
This Reality Lab owns one real-world evidence-transfer job: how a Primary 5/6 learner evaluates a claim such as “several laboratories got the same answer, therefore the answer is accurate” by separating interlaboratory agreement from accuracy and checking for common sources of bias.
It does not replace the broader eduKateSengkang owners for repeated measurements, measurement uncertainty, fair comparison, variables, method limitations, graph reading, conclusions or answer construction. Those pages teach the underlying scientific tools. Here we apply them to a communication object that appears in reports, product testing, environmental measurements, proficiency studies, standards work and scientific papers: the interlaboratory comparison.
Two Questions That Look Similar but Are Not the Same
| Question | What it examines | Useful evidence | What it cannot prove alone |
|---|---|---|---|
| Do laboratories agree? | Reproducibility or comparability among results under stated conditions | Results from multiple laboratories, spread, method details, repeated runs | That the shared answer is close to the true or reference value |
| Are the results accurate? | Closeness to an appropriate reference for the quantity | Reference value, certified reference material, validated reference method, traceable measurement chain where relevant | That every future sample or every laboratory will perform equally well |
| Do the laboratories disagree? | Differences that may arise from methods, instruments, operators, samples or random variation | Raw results plus comparison conditions | That the most different laboratory must be wrong |
The distinction is easier to see with a target board. Five arrows can land in a tight cluster away from the centre. The arrows are mutually consistent, but the cluster is displaced. In measurement language, a similar pattern can occur when a method is reproducible yet biased.
Do not stretch the target-board analogy too far. Real measurements involve reference values, uncertainty, sample preparation, instruments and models. The analogy simply protects one idea: closeness to each other and closeness to the right reference are different relationships.
Reality Lab Dossier: The Blue Reference Bottle
Consider an original composite case. A science network prepares a stable blue solution for a round-robin comparison. The organisers provide each laboratory with a sealed bottle from the same well-mixed batch. The best-supported reference value for substance Q is 50.0 mg/L under the stated measurement basis. Four laboratories receive bottles.
| Laboratory | Method family | Reported result | Three repeat readings inside laboratory |
|---|---|---|---|
| A | Method X | 52.0 mg/L | 51.9, 52.0, 52.1 |
| B | Method X | 52.1 mg/L | 52.0, 52.1, 52.2 |
| C | Method X | 51.9 mg/L | 51.8, 51.9, 52.0 |
| D | Method Y | 50.2 mg/L | 49.8, 50.2, 50.6 |
A, B and C agree extremely well. Their within-laboratory repeats are also tight. If you looked only at those three laboratories, “about 52 mg/L” would feel very secure. Laboratory D is less tightly repeated, yet its average lies much closer to the 50.0 mg/L reference value.
Now the investigation finds that Method X uses a conversion factor based on an assumed material property. The value copied into all three laboratories’ software is slightly too high for this particular material. The laboratories did not make three independent random mistakes. They inherited one shared assumption. That common step moved all three results in the same direction.
This constructed case mirrors a real measurement lesson. NIST has reported interlaboratory work in which some methods showed high repeatability and reproducibility while results could still be significantly biased because inaccurate material properties were used in calculations. Other methods could be less biased but less reproducible. The point for a Primary learner is not the advanced nanoparticle method. The transferable lesson is that different quality dimensions can move separately.
Observed, Claimed and Inferred
Return to Laboratories A, B and C.
- Observed: the reported results are 52.0, 52.1 and 51.9 mg/L.
- Observed: each laboratory’s three repeat readings are close.
- Observed: the laboratories used Method X.
- Claimed: the laboratories agree closely under these test conditions.
- Inferred too far: therefore 52.0 mg/L must be the accurate value.
- New evidence: an appropriate reference value is 50.0 mg/L and Method X shares a biased conversion factor.
- Better conclusion: the laboratories demonstrate strong agreement for Method X in this comparison, but their common result is biased high relative to the reference.
The scientific move is not to insult the laboratories. It is to describe exactly what the evidence supports.
The Five-Layer Agreement Check
When a report says “multiple laboratories agree,” inspect five layers before deciding how much that agreement means.
Layer 1: Were they actually testing comparable samples?
If Laboratory A receives a clear portion from the top of a bottle while Laboratory B receives sediment-rich material from the bottom, different results may describe different samples rather than different laboratory quality. Conversely, laboratories can appear to agree because all samples were prepared in the same way, even if that preparation altered the quantity of interest.
Ask about mixing, storage, transport, temperature, holding time, contamination, loss by evaporation, filtering or any treatment that can change the sample before measurement. Interlaboratory comparison begins before the instrument is switched on.
Layer 2: Were they measuring the same quantity on the same basis?
“50” is meaningless until we know 50 of what, per what, measured how. One laboratory might report mass concentration. Another might report number concentration. One might report a dry-mass basis while another uses wet mass. Agreement requires a shared measurement question, not merely similar-looking numbers.
Layer 3: How independent were the measurement routes?
Three laboratories can be independent organisations and still share critical dependencies. They may use the same instrument model, software library, calibration source, reference table, sample-preparation kit, conversion formula or training document. Shared components are not automatically bad; standardisation can improve comparability. But they matter when we interpret agreement as independent confirmation.
Layer 4: What reference anchors accuracy?
If no suitable reference exists, an interlaboratory comparison may still be useful. It can show how reproducibly different laboratories measure the object and reveal method differences. But it may be scientifically dishonest to convert that agreement into the stronger statement “the shared answer is definitely accurate.” Sometimes the best conclusion is narrower: the participating laboratories are mutually consistent within this comparison.
Layer 5: How far can the result travel?
A good result for one stable reference liquid does not guarantee identical performance for muddy river water, fruit juice, soil, smoke particles or another concentration range. Method performance can depend on the sample, range and purpose. The comparison supports what was actually tested first; extension needs evidence.
Representation Check: A Tight Cluster Can Be Visually Seductive
Suppose a report shows three dots at 51.9, 52.0 and 52.1. The horizontal axis runs from 51.5 to 52.5. The dots form a beautifully tight cluster. A reader may think “excellent accuracy” because the picture looks precise.
Now add a vertical reference line at 50.0 mg/L. It falls outside the displayed axis. Suddenly the graph’s persuasive message changes. The data did not change; the reference context changed.
For any agreement chart, ask:
- What does each point represent: one reading, a laboratory mean, or a corrected result?
- Is there an external reference or only a group average?
- Does the axis include the reference?
- Are uncertainty intervals shown?
- Are all laboratories using the same unit and reporting basis?
- Were any results excluded, and why?
A group average is useful for describing the group. It is not automatically a truth machine.
Consensus Value Is Not Automatically a Reference Value
Imagine 20 laboratories measure the same unknown sample and the average of their results is 75 units. Calling 75 the consensus value may be useful: it summarises what the group obtained. But if all 20 use the same biased method, the consensus can also be biased.
That does not make consensus useless. It changes the claim. Consensus can help identify outliers, compare laboratories and study reproducibility. Accuracy needs an appropriate connection to a reference that is justified for the measurement problem.
Worked Case 1: Three Thermometers Share One Bad Reference
Three school laboratories compare digital thermometers. Before the comparison, each thermometer was adjusted using the same reference device. Unknown to the schools, that reference device reads 0.8°C too high.
Later, all three thermometers read 25.8°C in a stable bath whose independently established temperature is close to 25.0°C.
- The three thermometers agree with one another.
- The agreement is not three independent confirmations of 25.8°C because the instruments share a calibration dependency.
- The independent bath reference reveals the common shift.
- The appropriate repair is to investigate the shared measurement chain, not to conclude that the bath must have changed because three devices agree.
Transfer: shared origin can create shared error. Count independent evidence by its dependency structure, not simply by the number of organisations or instruments.
Worked Case 2: The Odd Laboratory Might Be the Closer One
Four laboratories report 102, 102, 102 and 100 units. A student circles 100 and writes, “wrong because it does not agree with the others.”
Then a reference material certificate gives 100.0 units for the property on the same basis. The apparently odd result is now the closest of the four.
This does not prove the fourth laboratory is always better. It shows why majority agreement cannot replace the reference check. A scientific outlier is a result to investigate, not a result to condemn by vote.
Worked Case 3: Laboratories Disagree Because the Sample Changed in Transit
A composite environmental sample contains a volatile substance. Bottles are sent to three laboratories. Laboratory A receives and analyses its bottle the same day. Laboratory B tests after two days. Laboratory C tests after four days. Results fall in that order.
It would be easy to say, “The laboratories are not reproducible.” But the first question is whether the bottles still represented the same material by the time of testing. If the substance can be lost during storage, holding time becomes an alternative explanation for the disagreement.
Transfer: before diagnosing measurement disagreement, protect sample integrity. Comparison quality depends on both the laboratory and the object being compared.
Worked Case 4: Same Method, Different Operators
Five laboratories follow the same written procedure. Four report values between 20.0 and 20.3 units. One reports 23.8. Investigation shows that the fifth laboratory waited only 30 seconds for a colour reaction that the method specifies should develop for five minutes.
Here the between-laboratory comparison has done useful diagnostic work. It revealed that implementation differed even though the document name was the same. “Same method” on a form does not prove that every critical step was performed the same way.
But even after fixing the timing, a second question remains: does the correctly followed method measure the target quantity accurately? Reproducibility and accuracy still require separate evidence.
Worked Case 5: Different Methods, Similar Answer
Suppose Method X uses colour intensity and Method Y uses mass after a separation step. Independent laboratories using these different measurement principles obtain compatible results for a reference sample and for several ordinary samples.
This can strengthen confidence because the agreement is less likely to come from one shared method-specific assumption. Yet the conclusion should still match the evidence. If both methods fail on highly coloured samples, agreement on clear samples does not prove performance everywhere.
Transfer: agreement across genuinely different measurement routes can be particularly informative, but only for conditions that were actually tested.
Common Bias: The Hidden Bridge Connecting “Independent” Results
When several sources agree, ask what they share. Possible shared dependencies include:
- one reference material or calibrator;
- one conversion constant;
- one database of physical properties;
- one software package or algorithm;
- one sample-preparation procedure;
- one instrument family;
- one set of assumptions about density, purity, moisture or composition;
- one training protocol;
- one original dataset that later reports reanalyse.
Shared dependencies can be sensible and necessary. Scientific standards deliberately create common references so measurements can be comparable. The learner’s job is not to treat sharing as suspicious. It is to recognise that agreement created through a shared chain is not the same thing as several completely independent routes reaching the same conclusion.
Method Check: Repeatability, Reproducibility and Accuracy Are Different Jobs
These terms are often used casually, so keep the learner version simple.
| Idea | Learner-friendly question | Example |
|---|---|---|
| Repeatability | Does the same measurement set-up give similar results when repeated under closely similar conditions? | One laboratory measures the same stable sample several times. |
| Reproducibility | Do results remain reasonably compatible when relevant conditions change, such as laboratory, operator or equipment, according to the comparison design? | Several laboratories test distributed samples. |
| Accuracy | How close is the result to an appropriate reference for the quantity? | A laboratory result is compared with a suitable reference value. |
The definitions in advanced metrology are more technical than this classroom summary, and exact use can depend on context. For a Primary learner, the crucial protection is to stop using “consistent”, “reproducible” and “accurate” as interchangeable compliments.
What Evidence Would Strengthen the Claim That the Shared Result Is Accurate?
- A reference value suitable for the exact quantity and sample type.
- Results that agree with that reference within the justified measurement context.
- Independent methods based on different physical principles reaching compatible results.
- Evidence that reference materials and calibrations are appropriate and within scope.
- Stable, homogeneous comparison samples with documented handling.
- Results across more than one concentration or condition when the claim is intended to be broad.
- Transparent treatment of uncertainty and excluded data.
- Follow-up comparisons showing that an identified bias was corrected.
What Would Weaken a Strong Accuracy Claim?
- The only evidence is that several laboratories obtained similar numbers.
- All laboratories use the same unverified conversion or calibration source.
- The supposed reference is merely the average of the same laboratories being judged.
- Samples may have changed differently during transport or storage.
- Methods measure slightly different quantities or use different reporting bases.
- Only one easy sample was tested, but the conclusion claims all sample types.
- Large disagreement is hidden by reporting only a grand average.
- An outlying result is removed without a scientific reason.
The Grand-Average Trap
Suppose five laboratories report 48, 49, 50, 51 and 72. Their mean is 54. The sentence “the five-laboratory average was 54” is arithmetically correct. It is also a poor summary if the reader never sees the large disagreement.
An average can hide the shape of the evidence. Before accepting a group mean as the headline, inspect the individual results, the spread, the method groups and any justified reasons for differences. This is especially important when different methods form different clusters.
Do not turn this into a rule that averaging is bad. Averages are useful. The rule is narrower: never let a summary statistic erase evidence that matters to the scientific question.
Alternative Explanations for Laboratory Disagreement
If laboratories do not agree, there may be more than one plausible explanation. Healthy scepticism means generating possibilities and then seeking evidence that separates them.
- The sample was not homogeneous.
- The sample changed during storage or transport.
- Laboratories used different preparation steps.
- Instruments respond differently to interfering substances.
- One calculation used a different conversion basis.
- A transcription or unit error occurred.
- One result reflects ordinary random variation.
- A method works well in one range but poorly in another.
- A laboratory did not follow a critical timing or temperature condition.
- The reference value itself has uncertainty or is unsuitable for this sample.
The learner should not choose the most dramatic explanation first. Choose the explanation that best fits the evidence after checking the method and measurement chain.
How Far Can an Interlaboratory Result Travel?
A comparison can be strong evidence and still have boundaries.
- Across laboratories: the participating laboratories may not represent every laboratory.
- Across methods: agreement for Method X does not prove Method Y behaves the same way.
- Across sample types: clear water is not automatically equivalent to muddy water, food, soil or biological tissue.
- Across ranges: performance near 50 mg/L may not describe performance near the detection limit or at extremely high concentrations.
- Across time: instruments, reagents, software and staff can change.
- Across claims: accurate measurement of one property does not prove a product’s safety, effectiveness or overall quality.
The evidence should travel only as far as the comparison design can carry it.
Tempting but Invalid Reasoning
- “Most laboratories got 52, so 52 must be true.” Majority agreement can share a common bias.
- “The odd laboratory is wrong.” It may be wrong, or it may be closer to the reference. Investigate.
- “Three independent laboratories means three independent methods.” Organisations and measurement routes are different kinds of independence.
- “Tight repeats prove accuracy.” Tight repeats show consistency under those repeat conditions.
- “A reference value is exact truth with zero uncertainty.” Real reference values can have uncertainty and scope conditions.
- “If laboratories disagree, science has failed.” Disagreement can reveal hidden method, sample or measurement differences and help improve practice.
- “If laboratories agree once, the method is universally reliable.” One comparison has a defined sample, range, time and purpose.
Model and Measurement Limits
No measurement result arrives without a measurement model, even if the model is simple. A thermometer assumes a relationship between sensor response and temperature. A colour test assumes a relationship between colour intensity and concentration. A particle-count method may require assumptions about optical properties, density or shape. When laboratories share the same model assumptions, they can share the same model error.
At Primary level, you do not need advanced equations to reason correctly. Ask four plain questions:
- What was directly observed or measured?
- What calculation or conversion turned that observation into the reported quantity?
- What reference is used to judge accuracy?
- Which parts of that chain are shared by the laboratories?
Those questions are enough to stop “they agree” from becoming “therefore they are certainly right.”
PSLE-Style Transfer Case
Three laboratories are asked to determine the concentration of substance R in portions taken from one stable reference solution. The reference value is 80.0 units.
| Lab | Result | Method |
|---|---|---|
| P | 84.1 | M |
| Q | 84.0 | M |
| R | 83.9 | M |
Question 1: What can the results support directly?
Answer: The three laboratories obtained very similar results using Method M, so the comparison shows strong agreement among those laboratory results under the stated conditions.
Question 2: Do the results prove Method M is accurate for this reference solution?
Answer: No. The results are all around four units above the 80.0-unit reference. Their agreement does not remove the difference from the reference.
Question 3: Give one possible explanation that should be investigated.
Answer: The laboratories may share a calibration, conversion factor or method step that shifts all three results high. Other explanations are possible, so evidence is needed before deciding.
Question 4: A fourth laboratory using a different method reports 80.4 units. Is it automatically the best laboratory?
Answer: No. Its result is closer to the reference in this case, but one result cannot establish its overall performance. We would inspect repeat measurements, uncertainty, method suitability and performance on further samples.
Explained Practice 1: Agreeing Rain Gauges
Three electronic rain gauges at a test facility each report about 100 mm after a controlled water-delivery test. A calibrated collection vessel indicates that 95 mm was delivered to each gauge opening.
Best interpretation: the gauges agree with one another but appear to share a positive difference from the independent reference. Investigate shared calibration, geometry, conversion and test setup before calling 100 mm accurate.
Explained Practice 2: Disagreeing Soil Results
Three laboratories receive soil from the same field. Results differ greatly. The report later states that the soil was not mixed before subsamples were bottled.
Best interpretation: laboratory disagreement cannot be separated cleanly from sample heterogeneity. Different bottles may have contained different proportions of the material. A better comparison would first control the sample-distribution problem.
Explained Practice 3: Agreement Without a Reference
Six laboratories test a new property for which no accepted reference value is available. Five obtain values between 14.8 and 15.2; one obtains 18.1.
Best interpretation: five laboratories form a close cluster, which is evidence about comparability, while the 18.1 result deserves investigation. Without a suitable external reference, it is too strong to claim that 15.0 is proven accurate merely because it is the majority cluster.
A Four-Sentence Scientific Response Pattern
Do not memorise this as a compulsory exam template. Use it as a reasoning scaffold while learning.
- State the observed agreement: “The laboratories obtained closely similar values.”
- Name the limit: “This shows agreement but does not by itself establish accuracy.”
- Identify the needed check: “Compare the results with an appropriate reference and inspect shared method or calibration steps.”
- Conclude within scope: “Therefore the evidence supports reproducibility under these conditions, while accuracy requires the reference comparison.”
In an actual PSLE answer, use the evidence and wording demanded by the specific question. There is no universal magic phrase that earns marks independent of the science.
Delayed Independent Return
Return to this problem a day later without rereading the article first.
Draw two targets on paper. On Target A, draw four marks tightly grouped away from the centre. On Target B, draw four marks spread around the centre. Under each target, write what you can and cannot say about agreement and accuracy. Then add one sentence explaining why real laboratory evidence needs more than the drawing.
Next, invent a case where five laboratories agree because they share one conversion factor. Then invent a second case where laboratories disagree because the sample changed in transport. If you can distinguish those two mechanisms without prompts, the idea is beginning to transfer.
Useful eduKateSengkang Routes
- How to Decide Whether Two PSLE Science Measurements Are Different Enough for the Scale to Support a Real Difference
- How Repeating PSLE Science Investigations Affects Repeatability, Not Fairness
- How to Work Out What a PSLE Science Investigation Actually Measured
- How to Use Healthy Scepticism in PSLE Science Without Distrusting Every Result
- PSLE Science Reality Lab Vol No.047 — “Three Studies Agree” — Did They All Use the Same Dataset?
- PSLE Science Reality Lab Vol No.138 — “Method A Says 12, Method B Says 16” — Does One Method Have to Be Wrong?
Parent and Tutor Teaching Guide
Start with numbers simple enough that arithmetic does not steal attention from reasoning. Write three results on cards: 52.0, 52.1 and 51.9. Ask the learner what they notice. Most will correctly say the results agree. Then place a fourth card labelled Reference = 50.0. Ask what changed in the conclusion even though none of the original measurements changed.
In the second round, remove the reference card. Add a card saying, All three laboratories used the same conversion table. Ask whether this proves the result is wrong. It does not. It identifies a shared dependency that should be checked. This teaches the learner to distinguish reason for investigation from proof of error.
In the third round, provide four laboratory results where the odd result is closest to the reference. Ask the learner to explain why “majority wins” is not a scientific decision rule.
Finally, move the context. Use temperature, mass, water concentration and plant growth. Keep the evidence pattern the same while changing the surface details. The learner should eventually recognise the structure without being told that the lesson is “about accuracy”.
Authoritative Sources and Further Reading
- Singapore Examinations and Assessment Board — PSLE Science syllabus for examination from 2026
- Singapore Ministry of Education — Science Teaching & Learning Syllabus, Primary, 2023
- National Institute of Standards and Technology — VAMAS interlaboratory study on nanoparticle number concentration
- National Institute of Standards and Technology — Interlaboratory Comparisons
The NIST interlaboratory example is useful because it shows a real situation in which repeatability, reproducibility and bias did not all move together. The advanced measurement details are not required for PSLE Science; the evidence habit is.
The Quiet Return
Agreement deserves attention. When independent laboratories obtain compatible results, science gains useful information about how measurements behave across people, instruments and places. But agreement should not be promoted into a larger claim simply because the numbers look reassuring.
Ask what the laboratories shared. Ask whether the samples stayed comparable. Ask what quantity was measured. Ask where the reference comes from. Ask whether the claim is about consistency, reproducibility or accuracy.
Then make the smallest conclusion that the evidence can carry: these laboratories agree under these conditions, or, when the reference evidence supports it, these results also agree adequately with the reference for this purpose.
That is not weaker science. It is stronger reasoning because every word has an evidence job.