Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

Secondary 3 Mathematics Learning Guide | Histograms, Cumulative Frequency, Box Plots and Standard Deviation

Statistical diagrams are not decoration. They are compressed arguments about a distribution. A histogram shows how observations are distributed across intervals. A cumulative-frequency representation shows how many observations lie at or below a boundary. A box plot compresses positional information. Standard deviation describes spread around the mean.

This Secondary 3 Mathematics Learning Guide develops these ideas as one connected system. It deepens the broader Probability and Statistical Reasoning guide by concentrating on distribution shape, cumulative position and variability.

The current 2027 SEC G3 Mathematics syllabus listed by SEAB includes statistical diagrams, cumulative frequency, box-and-whisker plots and measures of spread. This guide keeps the statistical interpretation attached to every calculation.

Start With the Variable and the Question

Before drawing or reading a statistical display, identify the variable. Is it numerical or categorical? Is it discrete or measured continuously? Are we trying to compare centre, spread, shape, relative position or frequency?

The representation should serve the question. A cumulative-frequency curve is useful for percentile positions; a histogram is useful for interval distribution; a box plot is useful for comparing medians and interquartile ranges.

Histograms Represent Numerical Intervals

In a histogram, neighbouring bars touch because the horizontal axis represents adjacent numerical intervals rather than separate categories.

Under the equal-class-interval setting used in this course, bar height can represent frequency directly. The horizontal scale still matters: the bar covers a numerical interval, not a category name.

Worked Example 1: Read a Histogram Table

A data set is grouped into 0–10, 10–20, 20–30 and 30–40, with frequencies 4, 9, 11 and 6.

The modal interval is 20–30 because it has the highest frequency, 11. The total frequency is 4+9+11+6=30.

The histogram does not tell us the exact individual values inside each interval. It tells us how many observations occupy each range.

A Histogram Can Reveal Shape

A distribution may appear roughly symmetric, concentrated toward lower values, concentrated toward higher values, or show more than one cluster. These descriptions are qualitative and depend on the chosen intervals.

Changing class boundaries can alter the visual appearance. Therefore histogram interpretation should be tied to the actual grouped data, not to a vague impression of the picture alone.

Grouped Mean Uses Midpoints

When exact observations are unavailable, the mean of grouped data can be estimated by treating every observation in a class as though it were at the class midpoint.

Estimated mean = Σ(frequency×class midpoint)/Σfrequency.

Worked Example 2: Estimated Mean From Grouped Data

Using the previous intervals and frequencies, midpoints are 5,15,25,35.

Weighted total = 4(5)+9(15)+11(25)+6(35)=20+135+275+210=640.

Estimated mean = 640/30 ≈ 21.3.

It is an estimate because actual values within each interval are unknown.

Cumulative Frequency Is a Running Count

If interval frequencies are 4,9,11,6, the cumulative frequencies are 4,13,24,30.

The value 24 means that 24 observations lie below the upper boundary of the third interval. Each cumulative total contains all previous frequencies.

Worked Example 3: Locate the Median Position

There are 30 observations. The middle lies between the 15th and 16th ordered observations, or around cumulative position 15.5 for interpolation on a cumulative-frequency curve.

Since cumulative frequency is 13 by the end of 10–20 and 24 by the end of 20–30, the median lies inside the 20–30 interval.

A curve can estimate the corresponding numerical value within that interval.

Quartiles Divide Ordered Data Into Four Parts

The lower quartile Q1 marks roughly the 25th percentile, the median Q2 the 50th percentile, and the upper quartile Q3 the 75th percentile.

School conventions for exact position formulas can differ slightly depending on whether raw or grouped data are used. Follow the method required by the problem or diagram, but keep the interpretation stable: quartiles describe positions within the ordered distribution.

Percentiles Generalise Quartiles

The 70th percentile is a value at or below which about 70% of observations lie. It does not mean the score is 70% of the maximum possible score.

This distinction is important in examination and test contexts: percentile describes relative position within a group, not percentage marks.

Worked Example 4: Read a Percentile Position

In a data set of 80 observations, the 75th percentile corresponds to cumulative frequency approximately 0.75×80=60.

On a cumulative-frequency curve, move horizontally from cumulative frequency 60 to the curve, then vertically to the data axis to estimate Q3.

Interquartile Range Measures the Middle Half

IQR = Q3−Q1. It measures the width containing the middle 50% of ordered observations.

Because it ignores the outer quarters when measuring its width, it is less sensitive to extreme values than the full range.

Worked Example 5: Compare IQR

Group A has Q1=42 and Q3=58, so IQR=16. Group B has Q1=39 and Q3=64, so IQR=25.

Group B has a more spread-out middle half. That statement is about variability, not about which group has the higher median.

Box Plots Show a Five-Number Summary

A box plot shows minimum, Q1, median, Q3 and maximum. The box spans Q1 to Q3; a line inside marks the median; whiskers extend toward the extremes under the convention used by the problem.

Box plots are especially useful for side-by-side comparison because centre and spread can be inspected at the same time.

Worked Example 6: Compare Two Box Plots

Group A: minimum 20, Q1=35, median=50, Q3=61, maximum=76. Group B: minimum 30, Q1=38, median=54, Q3=60, maximum=67.

Group B has the higher median, 54 versus 50. Its IQR is 22, while A’s IQR is 26. Thus B has a slightly higher centre and a slightly tighter middle half.

However, this does not tell us every detail of the underlying distributions. Box plots deliberately compress information.

Range Uses the Extremes

Range = maximum−minimum. It is simple but depends entirely on the two extreme observations.

For Group A above, range=76−20=56. For Group B, range=67−30=37. Both the IQR and range suggest that A is more spread out, but they measure different parts of the distribution.

Standard Deviation Uses Every Observation

Standard deviation measures how dispersed values are around the mean. Small standard deviation means observations tend to lie closer to the mean; large standard deviation means greater spread.

Unlike range, standard deviation uses all observations. Unlike IQR, it is tied to the mean rather than quartile positions.

Worked Example 7: Same Mean, Different Standard Deviation

Data A: 48,49,50,51,52. Data B: 30,40,50,60,70. Both have mean 50.

For A, deviations from the mean are −2,−1,0,1,2. For B, they are −20,−10,0,10,20. Every B deviation is ten times the corresponding A deviation, so B’s standard deviation is ten times as large.

The equal means do not imply equal consistency.

Use the Calculator, But Interpret the Output

Approved calculators can calculate statistical summaries, but entering data correctly is only the first step. The student must know whether the problem asks for mean, median, standard deviation or another statistic and what that result means.

After calculator entry, check the number of observations, the approximate centre and the expected spread. A standard deviation larger than the full range, for example, should trigger immediate suspicion in ordinary finite data.

Comparing Distributions Requires More Than One Number

A stronger comparison usually comments on centre and spread together. For example: “Group B has a higher median but a smaller IQR, so its typical value is higher and its middle half is less dispersed.”

Avoid vague statements such as “B is better” unless the context explains what “better” means.

Worked Example 8: Contextual Comparison

Two delivery teams have mean delivery times 42 and 39 minutes, with standard deviations 4 and 9 minutes respectively.

Under this simplified comparison, Team B has the lower mean time but greater variability. Team A is slower on average but more consistent.

Whether speed or consistency matters more depends on the operational goal.

Misleading Statistical Displays

Truncated axes, inconsistent interval widths, omitted labels and selective time windows can alter visual impression. A graph can be technically constructed from real data and still encourage a misleading conclusion if the scale or framing is ignored.

Always read axes, units, interval definitions and sample size before interpreting shape or change.

Common Errors

Treating a histogram like a bar chart: remember that histogram bars represent numerical intervals and touch.

Reading percentile as percentage score: percentile describes position within an ordered group.

Comparing medians without spread: add IQR, range or standard deviation when variability matters.

Calling a grouped mean exact: it is an estimate based on class midpoints.

Independent Practice

1. Frequencies in intervals 0–10,10–20,20–30,30–40 are 5,8,12,5. Find total frequency and modal interval.
2. Find cumulative frequencies for Question 1.
3. Estimate the grouped mean using midpoints.
4. With 30 observations, identify approximately where the median position lies on a cumulative-frequency scale.
5. A data set has Q1=18, median=25, Q3=37. Find IQR.

6. Box plot A has median 50 and IQR 12. Box plot B has median 55 and IQR 20. Compare centre and middle-half spread.
7. Data sets have the same mean but one has much larger standard deviation. What does that imply?
8. A group has minimum 10 and maximum 46. Find range.
9. Explain why grouped mean is approximate.
10. Explain why a truncated vertical scale can exaggerate apparent change.

Explained Answers

1. Total=30; modal interval=20–30.
2. 5,13,25,30.
3. Midpoints 5,15,25,35 give weighted total 25+120+300+175=620; estimated mean=620/30≈20.7.
4. Around the 15th–16th ordered observations, or cumulative frequency about 15.5.
5. IQR=19.

6. B has higher median but greater IQR, so its centre is higher and its middle half is more spread out.
7. The data set with larger standard deviation is more dispersed around its mean.
8. 36.
9. Exact values inside each class are unknown, so class midpoints are used as representatives.
10. Starting the axis near the data values makes small numerical differences occupy a larger visual fraction of the displayed height.

Continue the Secondary 3 Learning Route

Continue with Angles, Parallel Lines, Polygons and Symmetry, Graphical Solutions, Intersections and Tangent Gradients, and Mathematical Reasoning, Proof and Communication.