Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

PSLE Science Reality Lab Vol No.125 | “95% Confidence Interval” — Does That Mean 95% of the Individual Results Must Fall Inside It?

PSLE-SCI-REALITY-0125

Wait, What? A 95% Confidence Interval Is Not a Box That Must Contain 95% of the Data Points

A graph shows an average value of 50 with a shaded band from 47 to 53. The caption says:

95% confidence interval: 47 to 53

A reader concludes, “That means 95% of all individual measurements must lie between 47 and 53.”

That is not what a confidence interval for a population mean usually means.

A confidence interval is an interval estimate for an unknown population quantity, such as a mean. Its confidence level describes the long-run performance of the interval-building method under its assumptions. It is not automatically a statement about what percentage of individual observations lie inside the interval.

The Reality Lab habit is: before interpreting a shaded band or ± value, ask what quantity the interval is estimating.

Quick Answer

  1. Find the quantity being estimated: a mean, difference, proportion, model parameter or something else.
  2. Do not assume the interval contains 95% of individual measurements.
  3. Do not automatically read a 95% interval as “95% probability this particular fixed interval contains the true value” in the ordinary frequentist interpretation.
  4. Check the interval width: a very wide interval signals limited precision.
  5. Check sample size and variability because they affect interval width.
  6. Check the assumptions used to build the interval.
  7. Keep the scientific conclusion inside the quantity the interval actually describes.

The Exact Learner Job This Page Owns

This page owns one real-world evidence-transfer job: interpreting a confidence interval shown in a scientific chart or report without confusing uncertainty about an estimated population quantity with the spread of individual observations.

It does not replace formal statistics instruction. It does not re-own averages, variation, error bars or statistical significance. Instead, it applies PSLE Science habits—identify the quantity, read the representation carefully, separate observation from inference, and respect model limits—to a common scientific communication object.

Original Reality Lab Case: The Seedling Heights

This is an original constructed teaching case.

A class measures the heights of 25 fictional seedlings grown under the same stated conditions. The sample mean is 50 mm. A statistical calculation gives a 95% confidence interval for the population mean of 47 to 53 mm.

The actual individual seedling heights range from 35 mm to 67 mm.

QuantityValue
Sample mean height50 mm
95% confidence interval for population mean47 to 53 mm
Observed individual range35 to 67 mm

There is no contradiction. The confidence interval is about uncertainty in the mean of the wider population, not a container designed to hold 95% of individual plants.

Observed, Estimated and Inferred

LayerStatement
ObservedThe 25 measured seedlings produced individual heights and a sample mean of 50 mm.
EstimatedThe statistical method gives an interval of 47 to 53 mm for the unknown population mean.
Unsupported shortcut“95% of seedlings must be between 47 and 53 mm.”
Better interpretationThe interval describes uncertainty in the estimate of the population mean under the method and assumptions used.

First Question: An Interval for What?

This is the most important question in the whole article.

An interval might be built for:

  • a population mean,
  • a difference between two means,
  • a proportion,
  • a regression coefficient,
  • a measured quantity with uncertainty,
  • or another estimated parameter.

Those intervals answer different questions. The phrase “95% confidence interval” is incomplete until you know which scientific quantity is being estimated.

Why Individual Results Can Sit Far Outside the Interval

Individuals can vary greatly even when the average of a large group is estimated quite precisely.

Imagine 1,000 leaves whose lengths vary from 4 cm to 12 cm. If you take a reasonably large representative sample, you may estimate the population mean to be around 8.0 cm with a fairly narrow confidence interval. That narrow interval does not say leaves themselves are all close to 8.0 cm. It says the mean may be estimated fairly precisely.

This distinction between variation among individuals and uncertainty in an estimated average is load-bearing scientific reasoning.

The Repeated-Sampling Meaning of 95%

NIST explains the frequentist interpretation this way: if the same population were sampled many times and a confidence interval were constructed by the same valid method each time, about 95% of those intervals would contain the true population parameter.

That is different from saying, after one interval has been calculated, “there is a 95% probability that this fixed interval contains the true value”. In the standard frequentist framework, the true parameter is treated as fixed; the interval-building procedure is what has the long-run coverage property.

For Primary 5/6 learners, the goal is not to master philosophical schools of statistics. The useful habit is simpler: do not casually turn a confidence level into a probability statement about individual observations or into a certainty score for a claim.

Confidence Interval Versus Prediction Interval

A confidence interval often estimates an unknown average or other parameter. A prediction interval is designed for a different job: estimating where a future individual observation may fall under the model.

Prediction intervals are usually wider because individual outcomes vary more than the uncertainty in the average. You do not need to calculate one at PSLE level. You only need to recognise that “where is the average?” and “where might one new individual result fall?” are different questions.

The Width Check: Narrow Is Not Automatically Better Science

A narrow confidence interval usually means the estimated quantity is more precise under the model. But precision is not the same as accuracy, representativeness or correctness.

A huge but biased sample can produce a narrow interval around the wrong target. A sensor with a systematic offset can give highly consistent data and a narrow interval. If the sampling design excludes an important group, more observations from the same biased source do not automatically repair the problem.

This is exactly why Reality Lab Vol No.117 separates small error bars from accuracy.

What Makes an Interval Wider or Narrower?

Without using formulas, a learner can understand three general influences:

  • More observations: often reduce uncertainty in an estimated mean when the sampling is appropriate.
  • More variability: usually makes the estimate less precise and the interval wider.
  • Higher confidence level: usually requires a wider interval because the method is trying to achieve higher long-run coverage.

These are relationships, not magic rules. The exact interval also depends on the model and assumptions.

The 90%, 95%, 99% Comparison

Suppose the same dataset gives:

Confidence levelIllustrative interval for the same mean
90%48 to 52
95%47 to 53
99%45 to 55

The exact numbers here are fictional, but the pattern illustrates a useful idea: demanding higher coverage generally widens the interval. You do not get more confidence for free.

The Sample-Size Check

A confidence interval from five observations and one from five thousand observations should not be read as though they carry the same information. With a small sample, a few unusual observations can move the estimate substantially. With a larger representative sample, the mean can often be estimated more precisely.

But “larger” does not automatically mean “representative”. If all 5,000 observations come from the wrong part of the system, sampling bias remains.

The Independence and Design Check

Many confidence-interval methods assume observations behave in particular ways—for example, that repeated measurements are sufficiently independent or that the statistical model describes the data reasonably well.

Taking 100 readings from the same unchanged object may not give the same scientific information as measuring 100 independently sampled objects. The calculation must match the evidence structure.

The Representation Check: Shaded Bands Can Hide Their Job

A graph may show a central line with a shaded band. But the band could represent a confidence interval, standard deviation, standard error, interquartile range, prediction interval or something else.

The shape alone does not tell you what the band means. The legend or caption must name it.

This is a valuable Reality Lab rule: never infer the meaning of an uncertainty graphic from its appearance alone.

Overlapping Confidence Intervals: Can You Decide Significance by Eye?

Not reliably in every situation. Whether two intervals overlap depends on what intervals were calculated and what statistical comparison is appropriate. Two 95% confidence intervals can overlap while a direct test of the difference is statistically significant, or fail to overlap under other conditions.

For a Primary learner, the safe rule is: do not invent a universal “overlap = same, no overlap = different” law. Use the analysis designed for the actual comparison.

What Would Strengthen a Confidence-Interval Claim?

  • The report states exactly which quantity the interval estimates.
  • The confidence level is stated.
  • The sample size and sampling method are transparent.
  • The interval-building method is appropriate for the data and study design.
  • The estimate and interval are shown together.
  • The public explanation distinguishes individual variability from uncertainty in the estimate.
  • Model assumptions and important limitations are disclosed.

What Would Weaken It?

  • A shaded band is shown without saying what it represents.
  • The caption says “95% of values lie here” when the interval is actually for a mean.
  • A narrow interval from a biased sample is presented as proof of accuracy.
  • The interval is calculated from non-independent repeated measurements but interpreted as population evidence.
  • The report hides a very wide interval and promotes only the central estimate.
  • The confidence level is turned into a “95% chance our claim is true” statement.

Worked Case 1: Narrow Mean Interval, Wide Individual Spread

A large sample of plant leaves has an average length of 8.0 cm with a 95% confidence interval of 7.9 to 8.1 cm. Individual leaves range from 4 cm to 12 cm. This is possible because the interval estimates the population mean, not the range of individual leaves.

Worked Case 2: Small Sample, Wide Interval

Five samples give an estimated mean of 30 units with a 95% interval from 18 to 42 units. A headline says “Average = 30”. The central estimate is correct, but the wide interval warns that the estimate is imprecise. Reporting only 30 hides important uncertainty.

Worked Case 3: Thousands of Biased Readings

A sensor is positioned only in the coolest corner of a large room and records 10,000 readings. The confidence interval for the mean at that sensor may be extremely narrow. That does not make it a precise estimate of the whole room’s average temperature. The sampling location is wrong for that broader claim.

Worked Case 4: Same Mean, Different Precision

Study A estimates a mean of 50 with a 95% interval of 49 to 51. Study B also estimates 50 but with an interval of 35 to 65. The central estimates match, but the evidence precision does not. Study B leaves much more uncertainty about the population mean.

Tempting Reasoning That Fails

  • “95% confidence interval means 95% of individual values fall inside.” Not when the interval is for a population parameter such as the mean.
  • “95% means a 95% probability this exact interval contains the fixed true mean.” That is not the standard frequentist interpretation.
  • “Narrow interval means accurate.” Precision can be high even when systematic bias remains.
  • “Wide interval means the study is useless.” A wide interval honestly communicates limited precision and may still rule out some possibilities.
  • “Two overlapping intervals prove no difference.” Visual overlap is not a universal statistical test.

Model and Measurement Limits

Confidence intervals rely on a statistical model and assumptions. If the sample is biased, observations are dependent in an unmodelled way, measurements are poor, or the interval formula is inappropriate, the stated coverage may not describe the real evidence well.

Also distinguish statistical confidence intervals from measurement uncertainty intervals. NIST metrology guidance uses carefully defined uncertainty concepts and sometimes reports expanded uncertainties with approximate confidence interpretations. A scientific report should tell the reader which kind of interval is being used.

How Far Can the Conclusion Travel?

A well-constructed 95% confidence interval can communicate how precisely an unknown population parameter has been estimated under the method’s assumptions. It cannot automatically describe 95% of individual outcomes, prove that the sample is representative, prove that the measurement is unbiased, or turn a scientific claim into 95% certainty.

PSLE-Style Transfer Case

A class measures the masses of 40 fruits and calculates an average of 120 g. A report gives a 95% confidence interval for the population mean of 116 g to 124 g. One fruit has a mass of 150 g.

Question: Does the 150 g fruit prove the confidence interval is wrong?

Reasoned answer: No. The interval is for the population mean, not for every individual fruit. Individual fruit masses can lie outside the interval while the interval still serves its intended job of estimating the mean.

Explained Practice

Practice A: A graph says “mean = 10, 95% CI 9.8–10.2”. Can you conclude almost all individual measurements are between 9.8 and 10.2? No.

Practice B: Two studies have the same mean but one has a much wider confidence interval. Which estimate is less precise? The one with the wider interval, assuming the intervals are comparable and valid.

Practice C: A study has a very narrow interval but sampled only one unusual location. What remains weak? Representativeness.

Delayed Independent Return: The C-I-R-C-L-E Check

  1. C — Confidence level: 90%, 95%, 99% or something else?
  2. I — Interval for what: mean, difference, proportion or another quantity?
  3. R — Range width: narrow or wide relative to the scientific question?
  4. C — Collection: was the sample representative and appropriate?
  5. L — Limits: what assumptions and measurement limits apply?
  6. E — Evidence boundary: what does the interval not tell you about individuals, causation or accuracy?

Parent and Tutor Teaching Guide

Use two drawings. In the first, draw many individual dots spread widely from left to right. Mark their average. In the second, imagine repeatedly taking samples from the same large population and calculating an interval around each sample mean. Ask: “Is the band trying to contain the dots, or estimate where the population mean is?”

Then show two intervals around the same central estimate—one narrow and one wide. Ask which leaves more uncertainty about the unknown average. This keeps the lesson visual and evidence-led without requiring formula memorisation.

Authoritative Sources

NIST’s statistics handbook explains the long-run coverage interpretation of confidence intervals and explicitly warns that a 95% confidence interval is not properly read as a 95% probability statement about a fixed interval containing a fixed true mean. That makes confidence intervals an excellent Reality Lab object: the mathematics is sophisticated, but the learner habit is familiar—identify what is measured, what is inferred and what the evidence cannot yet say.

The Quiet Return

An uncertainty band is only useful if you know what it is uncertain about.

Before reading a 95% interval, ask: 95% confidence in an interval for which scientific quantity?