A data set is not understood when one statistic has been calculated. The mean says something about centre. Quartiles and percentiles locate positions. Standard deviation describes spread. A cumulative-frequency graph reveals how observations accumulate. A careful comparison asks how these measures work together rather than treating one number as the whole story.
This fifty-third Secondary 4 Mathematics Learning Guide develops grouped-data interpretation as a connected statistical system. It belongs to the Secondary Mathematics Hub and S1–S4 Capability Map.
It deepens Cumulative Frequency, Box Plots and Standard Deviation, Mean, Median, Mode, Range and Comparing Data Sets and Histograms, Statistical Diagrams and Misleading Data.
The statistical question comes before the statistic
Before calculating, ask what the comparison is supposed to reveal. Is the issue typical performance, consistency, threshold attainment, relative position, spread or distribution shape? Different statistics answer different questions.
Centre tells us where the data sit. Spread tells us how widely they sit there.
Grouped data compress detail
When continuous values are placed into class intervals, the exact individual values are no longer visible. To estimate a grouped mean, we represent each class by its midpoint.
| Time t (min) | Frequency | Midpoint |
|---|---|---|
| 0≤t<10 | 4 | 5 |
| 10≤t<20 | 7 | 15 |
| 20≤t<30 | 9 | 25 |
| 30≤t<40 | 5 | 35 |
The midpoint is not being claimed as the true value of every observation. It is a modelling representative for that class.
Worked Example 1 | Estimate a grouped mean
Using the table above, total frequency is 25.
Estimated total of values:
4(5)+7(15)+9(25)+5(35)=525.
Estimated mean=525/25=21 minutes.
The word estimated matters because the exact observations inside each class are unknown.
Why midpoint estimates can differ from the true mean
If the 20≤t<30 class contains values mostly near 20, treating them all as 25 overestimates that class contribution. If they lie mostly near 30, the midpoint underestimates it. Grouping trades exact detail for a compact summary.
Standard deviation measures spread around the mean
Two groups can have the same mean and very different consistency. Standard deviation gives a numerical measure of spread: smaller standard deviation generally means values are more tightly clustered around the mean; larger standard deviation means greater dispersion.
Worked Example 2 | Same mean, different spread
Set A: 8, 9, 10, 11, 12.
Set B: 2, 6, 10, 14, 18.
Both sets have mean 10. But Set A is tightly clustered, while Set B is much more spread out. Therefore Set B has the larger standard deviation.
A comparison based only on the mean would miss this difference completely.
Grouped standard deviation also uses representative values
When standard deviation is calculated from grouped data, class midpoints again stand in for the unknown values within each interval. The result is therefore based on the grouped representation, not the original ungrouped data.
The calculator may perform the arithmetic, but the learner still needs to enter midpoints and frequencies correctly, distinguish population/statistical displays as required by the examination context, and interpret the result.
Quartiles divide ordered data into four parts
- Q1 marks the lower quarter.
- Q2 is the median.
- Q3 marks the upper quarter.
- Interquartile range IQR=Q3−Q1 measures the spread of the middle 50%.
Worked Example 3 | Interpret quartiles
A test distribution has Q1=54, median=68 and Q3=79.
- About 25% of observations are at or below the lower-quartile region.
- The median splits the ordered data into two halves.
- The middle 50% spans approximately 54 to 79.
- IQR=79−54=25.
Quartiles describe position in an ordered distribution. They do not tell us that every interval between them contains equal numerical width.
Percentiles generalise the same idea
The pth percentile is a value associated with approximately p% of the observations at or below that position in the ordered distribution. The 50th percentile is the median; the 25th and 75th percentiles correspond to the quartile positions.
Worked Example 4 | Read a percentile statement
A student’s score lies at the 80th percentile of a distribution. A careful interpretation is that the score is around the point with approximately 80% of observations at or below it, depending on the exact convention used in the question.
It does not mean the student scored 80% of the marks.
Cumulative frequency turns frequency into position
Cumulative frequency answers “how many observations have accumulated up to this boundary?” This makes it useful for estimating median, quartiles and percentiles from grouped data.
Worked Example 5 | Locate quartile positions
A grouped data set contains 80 observations.
- Q1 position is around 20th observation.
- Median position is around 40th observation.
- Q3 position is around 60th observation.
On a cumulative-frequency graph, read horizontally from the chosen cumulative frequency to the curve, then vertically to the data axis to estimate the value.
Worked Example 6 | 90th percentile from a cumulative total
If n=200, the 90th-percentile position is approximately 0.90×200=180. On the cumulative-frequency graph, locate cumulative frequency 180 and read the corresponding data value.
Comparing distributions needs at least two dimensions
A high-quality comparison commonly uses one statistic for centre and one for spread. For example:
Class A has a higher mean but also a larger standard deviation than Class B, so A performs higher on average but its results are less consistent.
This is more informative than declaring one class simply “better”.
Worked Example 7 | Mean and standard deviation comparison
| Mean | Standard deviation | |
|---|---|---|
| Group A | 72 | 5.1 |
| Group B | 69 | 8.4 |
Group A has the higher average and smaller spread. A careful conclusion is that A has higher central performance and greater consistency according to these measures.
We should not infer that every member of A scored higher than every member of B.
Worked Example 8 | Median and IQR comparison
| Median | IQR | |
|---|---|---|
| Machine X | 48 | 6 |
| Machine Y | 51 | 14 |
Y has the higher median output, but X has the smaller middle-50% spread. Therefore Y has the higher typical output by median, while X is more consistent by IQR.
Which centre should you use?
Mean uses all observations and is sensitive to extreme values. Median depends on order and is more resistant to extremes. In examination questions, use the measures given or requested, but understand what each can and cannot support.
Worked Example 9 | Extreme value effect
Data: 10, 11, 11, 12, 56.
Median=11. Mean=20.
The extreme value 56 pulls the mean upward strongly while leaving the median at the central ordered value.
A statistical statement should match the evidence
Statistics summarize. They do not automatically explain cause. A higher mean in one group does not prove a teaching method caused the difference unless the study design supports such a causal claim.
Similarly, a small standard deviation does not mean a group performed well; it means the values are tightly clustered. They could be consistently low.
Worked Example 10 | Consistently low
Group C has mean 42 and standard deviation 2. Group D has mean 68 and standard deviation 8.
C is more consistent, but D has much higher average performance. “More consistent” and “better average” are separate claims.
Grouped-data audit
- Is the data grouped or ungrouped?
- If grouped, what are the class midpoints?
- Does total frequency match the number of observations?
- Am I estimating the mean from representative class values?
- What measure describes centre?
- What measure describes spread?
- Where do quartile or percentile positions lie in cumulative frequency?
- Am I comparing at least one centre and one spread measure?
- Does my conclusion claim only what the statistics support?
Common failure modes
| Failure | Cause | Repair |
|---|---|---|
| Uses class boundary instead of midpoint for grouped mean | Representative value misunderstood | Use class midpoint |
| Calls grouped mean exact | Lost detail ignored | State that it is an estimate |
| Compares means only | Spread omitted | Add standard deviation or IQR comparison |
| Says smaller SD means higher results | Centre and spread confused | Interpret each statistic separately |
| Reads 80th percentile as 80% mark | Position and score confused | Interpret percentile as relative position |
| Overclaims every individual result | Summary statistic treated as full ordering | Limit claim to centre/spread evidence |
Independent practice
- Classes 0≤x<10, 10≤x<20, 20≤x<30 have frequencies 3, 5, 2. Estimate the mean.
- A data set has 120 observations. State the approximate cumulative-frequency positions of Q1, median and Q3.
- A distribution has Q1=24 and Q3=41. Find IQR.
- Group A has mean 76, SD 4.8. Group B has mean 73, SD 9.2. Compare carefully.
- Explain why grouped mean is generally an estimate.
- Explain why a smaller standard deviation does not necessarily mean higher performance.
Explained answers
1. Midpoints 5,15,25. Estimated total=3(5)+5(15)+2(25)=140. Total frequency=10. Mean=14.
2. Q1 around 30th, median around 60th, Q3 around 90th observation.
3. 41−24=17.
4. A has the higher mean and smaller standard deviation, so it has higher average performance and greater consistency by these measures.
5. Exact values inside each class are unknown; midpoints stand in for them.
6. Standard deviation measures spread, not the level of the centre.
Final thought
Statistical maturity is not knowing more formulas. It is knowing which summary answers which question, where grouping has removed detail, and how to make a comparison without claiming more than the evidence can support.
Compare centre with centre, spread with spread, and conclusions with the actual evidence.
Return to the Secondary Mathematics Hub.