Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

PSLE Science Reality Lab Vol No.299 | “Calibration Chart: R² = 0.99” — Is the Method 99% Accurate?

Series ID: PSLE-SCI-REALITY-0299

Wait, What? The Chart Says 0.99, and Your Brain Wants to Turn It Into 99%

A laboratory poster shows six reference samples, six dots and a straight line. In the corner is a neat label: R² = 0.99. A product brochure beside it says the instrument gives “highly accurate results”. The number looks so close to 1 that the tempting conclusion arrives almost automatically: R² of 0.99 means 99% accurate.

That conclusion is not justified. R² is a statistic about how well a fitted model describes variation in the data used for the fit. It is not a percentage accuracy score, not a promise that each reading is within 1% of the true value, and not a replacement for checking bias, precision, calibration range, residuals, reference values and independent verification.

This Reality Lab has one exact learner job: when a scientific chart displays R², decide what that number actually supports before letting it carry a stronger claim about measurement quality.

Quick Answer

  • R² is about fit. It describes how much of the variation in the fitted data is accounted for by the model, under the way the statistic was calculated.
  • Accuracy asks a different question. Are results close enough to appropriate reference values for the intended measurement job?
  • A high R² can coexist with bias. Points can line up beautifully while the entire relationship is shifted or scaled wrongly for the use being claimed.
  • Range matters. A good fit inside the calibrated interval does not automatically validate extrapolation beyond it.
  • Residuals and check standards matter. A single headline statistic can hide curvature, outliers or systematic error.
  • Do not convert 0.99 into “99% accurate”. Ask what was fitted, against what references, over what range, with what errors, and how independently the method was checked.

Owned Learner Job — and the Boundary

This article owns the real-world transfer problem created by a chart, report, instrument comparison or advertisement that displays an R² value and invites a learner to overread it as a quality percentage. It does not become the estate’s general owner of graph reading, correlation, calibration, uncertainty, fair testing or statistics. Those broader skills remain with their existing guides. Here, we apply them to one communication object: the impressive-looking R² badge.

The PSLE connection is reasoning, not terminology memorisation. The current 2026 PSLE Science assessment objectives include interpreting and analysing information, evaluating observations, information and methods, and communicating explanations and reasoning. The MOE Primary Science syllabus also explicitly values healthy scepticism: questioning observations, methods, processes and data. R² itself is not a magic PSLE keyword. The useful habit is learning not to let a familiar-looking number claim more than the evidence supports.

The Original Case: The Brilliant Sensor Chart

Imagine a fictional company called BrightLab Instruments. It is comparing a small light sensor with six certified reference light levels. The brochure shows this table:

Reference light levelSensor reading
1018
2028
3038
4048
5058
6068

The points lie perfectly on a straight line. A fitted line can have an R² of 1.00. Yet the raw sensor values are consistently 8 units above the references. If someone simply reports the sensor number as though it were already the reference quantity, the readings are biased even though the relationship is perfectly orderly.

This is the first shock worth keeping: a perfect fit does not automatically mean perfect agreement with the target values. Fit and accuracy are related only through the complete measurement method, including how the calibration equation is used and checked.

Observed, Calculated, Inferred, Claimed

LayerWhat belongs here?
ObservedThe reference values and corresponding instrument responses recorded during the test.
CalculatedA fitted equation, residuals and an R² value produced from those data.
InferredThe chosen model describes the pattern well enough for a stated purpose within a stated range, if other checks also support it.
Claimed“The method is 99% accurate everywhere.” This is a much stronger statement and needs different evidence.

When a scientific communication goes wrong, it often jumps from the calculated layer straight to a claim. Reality Lab readers rebuild the missing bridge.

What R² Actually Answers

In a common regression setting, R² is the coefficient of determination. NIST describes R-squared as a statistic connected to lack of fit of a predictive model and notes that a perfectly fitting model yields R² = 1. The important phrase is fitting model. The number is about the relationship between data and the fitted model; it is not labelled “accuracy percentage”.

Suppose reference concentration increases and instrument signal increases almost perfectly with it. A high R² tells us the fitted relationship captures the pattern very well. That can be useful evidence for a calibration. But a trustworthy calibration also needs the correct model, suitable standards, an appropriate working range, acceptable residual behaviour, sufficient precision and independent checks. NIST calibration guidance explicitly discusses bias, precision, outliers, model validation and check standards as separate issues.

The 0.99 Trap: Why It Looks Like a Percentage

Humans are trained from school to see decimals between 0 and 1 as percentages. 0.50 becomes 50%; 0.99 becomes 99%. Mathematically, you can express 0.99 as 99 hundredths. But that does not change the meaning of the quantity.

A football team winning 0.99 of its games, a solution containing a mass fraction of 0.99, and a regression with R² = 0.99 are not the same measurement job. The number alone does not tell you the noun.

Reality Lab habit: before converting a number, name the quantity.

Case File 1 — A Straight Line With a Wrong Offset

Reference temperatures are 10°C, 20°C, 30°C, 40°C and 50°C. A fictional sensor reads 13°C, 23°C, 33°C, 43°C and 53°C. The relationship is perfectly linear. The sensor response changes by the right amount every time, but it reads 3°C high before correction.

If the calibration equation correctly learns and removes that offset, the sensor may still become useful. If the brochure simply reports “R² = 1.00, therefore 100% accurate,” the reasoning fails. The fit does not erase the need to handle the offset.

Case File 2 — A Beautiful Fit Over the Wrong Range

A water-quality instrument is calibrated from 0 to 50 units. The points fit very well. Someone then uses the equation to report a value corresponding to 140 units and says, “The calibration has R² = 0.999, so the result must be reliable.”

The missing question is whether the method was shown to work at 140. A calibration can be excellent over its validated interval and unreliable outside it. Detector response can curve, saturate, change sensitivity or become noisy. FDA guidance gives the same broad warning in a different context: a concentration far outside the range of the standards used to build the calibration is not supported merely because the standard curve exists.

Case File 3 — High R², Curved Residuals

Imagine 40 points spanning a large range. A straight line looks excellent overall and gives R² = 0.995. Yet when you calculate each residual—the difference between a measured point and the fitted line—the low values are mostly above the line, middle values mostly below, and high values above again.

That curved residual pattern is a clue that a straight-line model may be missing structure. A headline R² can remain high because the overall trend is strong. This is why model checking does not end with one attractive statistic.

Case File 4 — Two Methods With the Same R²

Method A and Method B both report R² = 0.99 against reference samples. Method A has residual errors mostly within ±0.2 units. Method B has residual errors mostly within ±5 units because its working range is much wider. Which is “more accurate”?

You cannot answer from R² alone. You need the actual error sizes, units, range and intended use. Equal R² values do not mean equal uncertainty or equal suitability for a decision.

Case File 5 — The Independent Check Standard

A student team builds a calibration with five standards, fits a line and gets R² = 0.998. Instead of celebrating immediately, they test a sixth reference sample that was not used to build the line. The reference is 35 units; the calibrated method reports 35.3. They repeat this check on several days and track the differences.

Now the evidence is stronger. The fitted data say the relationship is orderly, while the check standard tests how the completed calibration behaves on independent known material. NIST explicitly describes check standards as a mechanism for estimating uncertainty associated with calibrated values. The extra evidence answers a question R² does not answer by itself.

Representation Check: What Did the Chart Choose to Show?

  • Are the actual calibration points visible, or only the fitted line?
  • Are axes labelled with quantities and units?
  • Is the calibration range shown?
  • Are residuals or errors reported?
  • Is R² rounded so heavily that meaningful differences disappear?
  • Does the chart show repeated measurements or just one point at each reference?
  • Are outliers visible, explained and handled transparently?
  • Does a large range make smaller but important errors look tiny on the graph?

A chart can be scientifically honest and still leave important information off the first panel. The learner’s job is not to accuse the chart of deception. It is to ask what evidence is present and what evidence would be needed for the stronger claim.

Comparison and Baseline Check

If two instruments advertise R² values of 0.998 and 0.995, it is tempting to rank the first as “more accurate”. That can fail for several reasons. The instruments may have been tested over different ranges, against different references, with different numbers of points, under different environments or using different models. Even if the R² values were calculated comparably, the difference may say less about practical performance than bias, repeatability or uncertainty does.

Before comparing badges, align the scientific object being compared.

CheckQuestion
ReferenceWere both methods compared with suitable known values?
RangeWas the same measurement interval tested?
ReplicatesWere repeated observations available?
ModelWas the same type of fitted relationship used?
ErrorsHow large were residuals and independent check errors?
ConditionsWere temperature, sample matrix, instrument settings and preparation comparable?

Method and Variable Check

R² can be affected by the spread of the x-values as well as by the model and errors. If standards span a very wide range, the overall trend can dominate the calculation. A method can therefore show a high R² while still making errors that matter near a decision threshold.

Suppose a safety-independent fictional quality limit is 10.0 units. The method reports values with errors of ±2 units near that limit. A chart spanning 0 to 1,000 units may still have a spectacular R². Yet the uncertainty around 10 could matter for the actual decision. Scientific usefulness is job-dependent.

Alternative Explanations for an Impressive R²

  • The method truly has a strong, stable relationship with the reference quantity.
  • The test covered a very broad range, making the general trend dominate smaller errors.
  • A few influential points at the ends of the range strongly constrain the line.
  • The chosen model fits the calibration data but has not been checked independently.
  • Systematic bias moves all results together without destroying the orderly relationship.
  • Repeated points are not independent and make the chart look more populated than the evidence base really is.
  • A nonlinear relationship happens to look linear within the tested interval.

Alternative explanations are not excuses to distrust every good fit. They are reasons to gather the right next evidence.

What Evidence Would Strengthen the Claim?

  • Reference materials with appropriate known values and uncertainty.
  • Calibration points covering the intended working interval.
  • Repeated measurements showing acceptable precision.
  • Residual plots without obvious unmodelled structure.
  • Independent check standards not used to fit the model.
  • Bias or recovery information showing closeness to reference values.
  • Tests across relevant days, operators, instruments or sample types when the use requires them.
  • Clear rules for what happens outside the calibrated range.

What Evidence Would Weaken the Claim?

  • Only the fitted line and R² are reported; raw points are hidden.
  • The intended samples fall outside the calibration range.
  • Check standards show systematic error.
  • Residuals show a curve or changing spread.
  • The model works for pure standards but not the real sample matrix.
  • Replicate results vary far more than the chart suggests.
  • The brochure translates R² directly into an accuracy percentage without supporting evidence.

How Far Can the Conclusion Travel?

A careful statement might be:

“Within the tested calibration range, the fitted model described the reference-response data very closely, with R² = 0.99. This supports a strong fit, but measurement accuracy must be judged using the actual errors, reference checks, precision, uncertainty and intended conditions.”

Notice the boundary. The statement does not say “99% accurate,” does not claim every future reading is correct, and does not extend beyond the tested range.

Tempting Reasoning That Fails

Tempting sentenceWhy it failsRepair
R² = 0.99, so it is 99% accurate.R² is not an accuracy percentage.State that the model fits the fitted data strongly; check errors separately.
R² is almost 1, so there is no bias.A systematic offset or scale error can coexist with an orderly relationship.Compare calibrated results with references and inspect bias.
The line is straight, so it works at any concentration.The validated range is finite.Keep predictions within the tested range unless new evidence validates extension.
Method A has the higher R², so it is always better.Different ranges, errors and use cases can make R² values incomparable.Align the comparison and examine performance measures relevant to the job.

Model and Measurement Limits

Every model compresses reality. A calibration equation may treat one interval as linear even though detector physics becomes nonlinear outside it. Reference values have uncertainty. Sample preparation can introduce variability. Temperature can shift response. Real samples can contain interfering substances absent from clean standards. R² does not absorb all of these into one universal truth number.

That is not a weakness of science. It is how science becomes more trustworthy: each quantity has a job, and several quantities can be combined to support a stronger conclusion.

PSLE-Style Transfer Case — The Colour Sensor

A student tests a colour sensor using five known dye concentrations. The graph of sensor signal against concentration gives R² = 0.997. A sixth known sample has a concentration of 25 units. The calibration estimates 29 units.

  1. What does the R² value support? The fitted model closely describes the pattern in the calibration data.
  2. Does it prove the method is 99.7% accurate? No.
  3. What new observation matters? The independent sample is estimated 4 units high.
  4. What should the student investigate? Bias, sample preparation, calibration model, instrument condition and whether the sixth sample lies inside the valid range.
  5. What is a bounded conclusion? The calibration points fit the model strongly, but the independent check shows that fit alone is insufficient to establish acceptable measurement performance.

Practice Laboratory — Four Mini Claims

Practice A

A graph reports R² = 1.000 using only three calibration points. Is the method proven perfect?

Answer: No. The three points fit the chosen model perfectly, but the method still needs appropriate range, repeated measurements, reference checks and independent validation for its intended use.

Practice B

Two sensors both have R² = 0.999. One has an average reference error of 0.1 unit; the other, 8 units. Are they equally accurate?

Answer: Not on the information given. The equal R² values describe fit, while the error evidence differs greatly.

Practice C

A calibration works from 0 to 100 units. The unknown produces a signal corresponding to 180. What is the first evidence question?

Answer: Whether the method has been validated beyond 100. A high R² within 0–100 does not automatically justify extrapolation to 180.

Practice D

A brochure shows only a line and R² = 0.998 but no points. What would strengthen your evaluation?

Answer: The raw calibration points, number of standards, range, residuals, repeatability, reference information and independent check results.

Delayed Independent Return

Come back later without looking above. Explain these three statements in your own words:

  1. Why can R² = 1 coexist with a constant measurement offset?
  2. Why does a good calibration fit not automatically validate an out-of-range sample?
  3. Which evidence would you want in addition to R² before accepting an accuracy claim?

If you can answer all three with the quantity names intact—fit, reference error, range and independent check—you have transferred the habit.

Route to the Existing PSLE Science Owners

Reality Lab applies the skills; it does not steal them. For repeated measurements and disagreement, continue to How to Read Repeated PSLE Science Results When the Measurements Do Not Match Exactly. For instrument checks against known values, use How to Use a Reference Value to Check a PSLE Science Measuring Instrument Before Trusting Its Readings. For anomalies, use How to Handle an Anomalous PSLE Science Result Without Deleting It Just Because It Looks Wrong.

Parent and Tutor Teaching Guide

Do not teach this by making a Primary 5/6 learner memorise the formal equation for R². The transferable lesson is simpler and more powerful: one impressive number cannot silently change its scientific job.

  1. Show a fictional calibration chart labelled R² = 0.99.
  2. Ask the learner to name what was directly measured.
  3. Ask what was calculated from those measurements.
  4. Offer the false sentence “therefore 99% accurate” and ask what extra evidence would be required.
  5. Give an independent reference sample that the calibration misses.
  6. Return two days later with a different context—a force sensor, light sensor or colour meter—and ask the learner to reconstruct the reasoning without the original wording.

The goal is not suspicion. It is disciplined interpretation. A learner should be able to say, calmly, “That R² is useful evidence about fit. Now show me the evidence about measurement error and range.”

Authoritative Sources

The Quiet Return

R² can be a very useful number. The mistake is not using it. The mistake is asking it to answer a question it was never designed to answer.

When a chart says R² = 0.99, do not hear “99% accurate.” Hear a better question:

What fits what—and what separate evidence shows that the completed measurement is good enough for the claim?