PSLE-SCI-REALITY-0138
Wait, What? Two Careful Scientists Can Measure the Same Thing and Still Report Different Numbers
Imagine that two laboratories receive portions of the same material. Laboratory A reports 12 units. Laboratory B reports 16 units.
A quick reaction is tempting: “One of them must be wrong.”
Sometimes one result really is wrong. But science does not allow us to jump there just because two numbers differ. Different methods can respond differently to the same sample. They can use different preparation steps, different physical signals, different calibration models, different assumptions and different ways of handling interfering material. Each may be measuring a slightly different version of the scientific quantity we thought was one simple thing.
The stronger question is: why do the methods disagree, and which method is fit for the exact claim we want to make?
Quick Answer
- Check that both methods are truly trying to measure the same quantity on the same basis.
- Trace the sample preparation used before each measurement.
- Ask what physical signal each method observes directly and what it calculates or infers afterward.
- Compare calibration range, uncertainty, interferences and known method bias.
- Do not average the two numbers simply to make the disagreement disappear.
- Use reference materials, controlled comparisons or an independent third method when the scientific decision requires stronger evidence.
The Exact Learner Job This Page Owns
This Reality Lab owns one real-world transfer job: evaluating a report, product comparison, research figure or laboratory claim when two legitimate measurement methods give different answers for what appears to be the same scientific quantity.
It does not replace the canonical owners for measurement uncertainty, accuracy and precision, calibration, variables, sampling or alternative explanations. It applies those ideas to a communication object that appears often in real science: method disagreement.
- How to Reconcile Two Pieces of PSLE Science Evidence That Seem to Disagree
- How Scientific Evidence Works | From Observation to a Claim You Can Defend
- How to Tell a PSLE Science Method Limitation From a Mistake in the Investigation
- Reality Lab Vol No.101: Were the Two Reports Measuring the Same Thing With the Same Test?
Original Reality Lab Case: The Same Powder, Two Methods
This is an original composite teaching case using fictional Substance R and constructed data.
A powdered material contains Substance R. The material is mixed carefully and divided into equal portions.
| Method | Preparation | Direct signal | Reported result |
|---|---|---|---|
| Method A | R is extracted into a liquid | Light absorbed by the extract | 12 units |
| Method B | Powder measured more directly | Instrument signal from the bulk material | 16 units |
Does Method B prove that Method A is wrong? Not yet. Perhaps Method A does not extract all of Substance R. Perhaps Method B also responds to another substance. Perhaps the methods define or calculate the target differently. Perhaps one calibration is outside its strongest range. Perhaps the powder was not as uniform as expected. Perhaps 12 and 16 are both compatible once measurement uncertainty is considered. Or perhaps one method really does contain an error.
The difference is a scientific clue, not a verdict.
First Check: Are the Two Methods Measuring the Same Quantity?
Scientific words can hide different measurement definitions. “Moisture”, “particle size”, “brightness”, “concentration”, “surface temperature” and “strength” can each be operationally defined in more than one way.
Two instruments may display the same unit while responding to different physical properties. One method may measure a surface. Another may average through a volume. One may measure a dissolved portion. Another may include particles. One may report a direct signal. Another may use a model to estimate the quantity from several signals.
Before deciding which answer is right, reconstruct the measurement definition.
Observed, Inferred and Calculated
| Layer | Method A | Method B |
|---|---|---|
| Direct observation | Instrument observes light response | Instrument observes another physical signal |
| Preparation | Target is extracted first | Bulk material is measured |
| Inference/model | Signal converted to amount using calibration | Signal converted to amount using a different model |
| Reported result | 12 units | 16 units |
The numbers at the bottom look comparable. The routes that created them are not identical. Good scientific reading follows the route upward before deciding what the difference means.
Second Check: Did Preparation Change What Could Be Measured?
Preparation is part of the measurement. Filtering, grinding, dissolving, drying, extracting, heating or separating can change which portion of a sample reaches the instrument.
If Method A extracts only 80–90% of the target under some conditions, a lower result might arise even when its instrument performs perfectly. If Method B measures the whole sample but has interference from another component, its result might be higher.
This is why “same sample” does not automatically mean “same measurement pathway”.
Third Check: Could Another Substance Affect One Method?
Methods differ in selectivity. A sensor or instrument may respond strongly to the target but also respond weakly to other materials. The surrounding sample can therefore matter.
A method that works beautifully in clean water may behave differently in muddy water. A colour measurement can be influenced by naturally coloured substances. A physical sensor can respond to temperature as well as the intended quantity. A camera-based measurement can depend on illumination.
The learner does not need specialist analytical chemistry. The transferable idea is enough: ask what else can move the signal.
Fourth Check: Were Both Methods Calibrated Over the Relevant Range?
A method can be excellent inside the range where it has been tested and less trustworthy outside it. If the sample produces a signal near the edge of calibration, one method may need dilution, a different range or a different model.
Reality Lab Vol No.095 already owns the specific job of reading beyond a calibration curve. Here, calibration is only one possible explanation for why two methods disagree.
Fifth Check: How Large Is the Measurement Uncertainty?
Suppose Method A reports 12 ± 3 units and Method B reports 16 ± 3 units under an appropriate uncertainty description. The two central values differ, but the uncertainty ranges overlap strongly. The disagreement may be less dramatic than “12 versus 16” first appears.
Now suppose Method A reports 12 ± 0.2 and Method B reports 16 ± 0.2. The difference deserves much more investigation.
Uncertainty does not automatically tell us which method is correct. It tells us how much spread or doubt belongs around the reported result under the stated measurement model.
NIST’s Useful Lesson: Different Methods Can Constrain Different Parts of the Problem
The U.S. National Institute of Standards and Technology has discussed situations in which different measurement methods provide different answers and how combining or comparing methods can reveal the assumptions and uncertainties of each. The important learner lesson is not that methods should always be averaged. It is that disagreement can expose which parameters or assumptions each method handles well.
A scientific method is not a camera taking a perfect picture of reality. It is a structured interaction between the object, preparation, instrument, calibration, model and uncertainty.
Should We Just Average 12 and 16 to Get 14?
Not automatically.
An average is justified only when there is a scientific reason to combine the results. If Method A has a known low bias in this sample type and Method B has a different interference, the number 14 may not be a better answer at all. If the two methods estimate slightly different quantities, averaging can create a number with no clear physical meaning.
First understand the disagreement. Then decide whether combination makes sense.
What Evidence Would Strengthen the Comparison?
- A reference material with a well-characterised value measured by both methods.
- Several samples spanning low, medium and high values rather than one sample.
- Replicate measurements showing each method’s repeatability.
- Known additions or controls that test recovery and interference.
- Clear documentation of sample preparation and calibration.
- Uncertainty statements appropriate to each method.
- A third independent method whose measurement pathway differs meaningfully from the first two.
What Would Weaken the Claim That “Method B Proves Method A Wrong”?
- Only one sample was compared.
- The two methods use different definitions of the target quantity.
- One method measures a filtered or extracted fraction while the other measures a whole sample.
- Known interferences differ between methods.
- Uncertainty is large compared with the difference.
- No suitable reference material or independent comparison exists.
- The report highlights whichever method gives the more dramatic marketing result.
Worked Case 1: Surface Temperature Versus Internal Temperature
An infrared instrument reads 38°C at a surface. A probe inside the object reads 33°C. Is one thermometer wrong? Not necessarily. They are sampling different locations and may be measuring different thermal states. The first question is whether the intended claim concerns the surface or the interior.
Worked Case 2: Filtered Water Versus Whole Water
Method A filters water before measuring a target. Method B measures a preparation that includes material associated with particles. Their results differ. The disagreement may arise because the operational definitions are different. “Total” and “dissolved” forms are not interchangeable simply because both reports use the same substance name.
Worked Case 3: Two Cameras Measure the Same Leaf Colour
Camera A reports an index value of 42 and Camera B reports 50. Before choosing a winner, check lighting, white balance, wavelength sensitivity, calibration reference and image processing. A digital number can be partly a property of the measurement system.
Worked Case 4: One Method Is More Precise but More Biased
Method A repeatedly gives 12.0, 12.1 and 12.0 units. Method B gives 15.4, 16.1 and 15.8 units. If a trusted reference is around 16, Method A may be highly repeatable yet systematically low. Close repetition is not enough to prove closeness to the target value.
Tempting Reasoning That Fails
- “The newer machine must be right.” Age does not determine fitness for purpose.
- “The more expensive method must be right.” Cost is not evidence.
- “The method with more decimal places is better.” Displayed digits do not create measurement quality.
- “Whichever result agrees with our expectation is correct.” That is confirmation bias, not method evaluation.
- “Two methods disagree, so science is unreliable.” Disagreement is often how limitations become visible and methods improve.
- “Average them and the problem disappears.” Arithmetic cannot replace scientific diagnosis.
Model and Measurement Limits
No measurement method is simply “good” or “bad” in every situation. A method can be excellent for one material, range and decision while unsuitable for another. It can be precise but biased, selective but slow, broad but less sensitive, or direct in one sense but model-dependent in another.
This is why metrology uses ideas such as traceability, uncertainty, validation and interlaboratory comparison. Those concepts help scientists state what a result means, how it connects to references and how strongly it can support a decision.
How Far Can the Conclusion Travel?
If two methods disagree on one material, we can say that the measurement system deserves investigation. We cannot immediately say that one method is universally wrong, that every sample will show the same difference, or that the higher number is automatically the more truthful one.
If repeated comparisons across appropriate reference materials show a consistent method-specific bias, the conclusion can become stronger and more specific.
PSLE-Style Transfer Case
Two groups measure the amount of water in the same kind of material. Group A dries the material and uses mass change. Group B uses a sensor whose signal depends on water content and calibration. Group A reports 18%, while Group B reports 22%.
Question: Why is it not scientifically valid to say immediately that Group B’s sensor is wrong?
Reasoned answer: The two methods use different measurement pathways and may have different calibration, preparation, uncertainty and interference effects. Their disagreement must be investigated using suitable controls or references before deciding which result is inaccurate.
Explained Practice
Practice A: Two methods disagree only for muddy samples but agree for clean water. What clue does that give? The sample matrix may affect one or both methods.
Practice B: Two methods differ by 0.3 units, but each has uncertainty of about ±1 unit. Should the central-value difference be treated as decisive? No. The uncertainty must be considered.
Practice C: A trusted reference material is measured as 10.1 by Method A and 13.8 by Method B when the reference value is near 10. Which method gains support in this test? Method A. That still does not prove Method A is superior for every sample and condition.
Delayed Independent Return: The M-E-T-H-O-D Check
- M — Measurand: Are both methods really measuring the same defined quantity?
- E — Evidence path: What does each instrument observe directly?
- T — Treatment: How was the sample prepared?
- H — Hidden influences: What interference or bias could affect each method?
- O — Operating range: Are calibration and uncertainty suitable here?
- D — Decision: What extra evidence would show which result is fit for the intended claim?
Parent and Tutor Teaching Guide
Use a familiar object such as temperature. Ask a learner why a forehead thermometer, room thermometer and food probe can show different values without any of them necessarily being defective. Then shift from location differences to method differences: two tools may interact with the same system in different ways.
Next, give two fictional laboratory numbers and ask the learner to generate at least three possible explanations before choosing one. The goal is to protect scientific reasoning from the false binary “A is right, therefore B is wrong” when the evidence has not yet earned that conclusion.
Authoritative Sources
- Singapore Examinations and Assessment Board — 2026 PSLE Science Syllabus
- Ministry of Education, Singapore — 2023 Primary Science Teaching and Learning Syllabus
- NIST — What To Do When Measurement Methods Produce Different Answers
- NIST — Metrological Traceability: Frequently Asked Questions and Policy
The current PSLE Science frame explicitly includes interpreting and analysing information, evaluating observations, information and methods, and communicating reasoning. The 2023 Primary Science syllabus also encourages healthy scepticism, consideration of uncertainty and more than one possible explanation. Method disagreement is an ideal real-world application of those habits.
The Quiet Return
Different numbers do not tell you which scientist failed.
They tell you where to look more closely at the path from reality to measurement.
When two methods disagree, do not choose a winner first. Reconstruct what each method actually did.