Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

Secondary 4 Mathematics Learning Guide | Histograms, Statistical Diagrams and Misleading Data

A statistical diagram can be numerically correct and still create a misleading impression. Scale, grouping, omitted context and visual design affect what a reader notices first. Secondary 4 data work therefore requires more than reading bars or plotting points: the learner must ask what the display represents, what it hides and whether the conclusion is supported.

This twenty-fourth Secondary 4 Mathematics Learning Guide develops statistical diagrams and visual evidence. It belongs to the Secondary Mathematics Hub and S1–S4 Capability Map. It focuses on bar charts, line graphs, pie charts, dot plots, stem-and-leaf diagrams, histograms with equal class intervals, cumulative-frequency displays, box plots and common forms of misleading presentation.

Current syllabus connection: upper-secondary Statistics and Probability includes interpreting and representing data in tables and several diagram types, together with measures of centre and spread and critical interpretation of representations. The examples here are original teaching material.

Choose a diagram because of the data structure

DisplayUseful forWhat to watch
Bar chartSeparate categoriesBars normally have gaps; category order may be arbitrary
Line graphChange across an ordered variable such as timeConnecting points implies an ordered progression
Pie chartParts of one wholeAngles and sectors must represent proportions of the same total
Dot plotSmall numerical data setsRepeated values and clustering remain visible
Stem-and-leafOrdered numerical data while preserving actual valuesKey and place value must be stated clearly
HistogramGrouped continuous numerical dataClass intervals touch; this guide uses equal class widths
Cumulative-frequency graphThreshold counts, median, quartiles and percentilesRead from cumulative total to variable carefully
Box plotComparing median and spread compactlyIt hides individual data values

Bar chart and histogram are not interchangeable

A bar chart usually represents categories such as transport type or favourite subject. The bars are separate because the categories are distinct.

A histogram represents grouped numerical intervals. Adjacent classes such as 0≤t<10 and 10≤t<20 touch because the numerical scale is continuous across the boundary.

For the equal-class-width histograms used in this guide, bar heights can directly represent frequencies when the class widths are the same. If class widths are unequal, height alone cannot be interpreted the same way without additional treatment.

Worked Example 1 | Read an equal-width histogram

A histogram groups travel times into 0–10, 10–20, 20–30 and 30–40 minutes, with frequencies 4, 12, 16 and 8.

The modal class is 20–30 minutes because it has the greatest frequency, 16.

Total frequency = 4+12+16+8=40.

The proportion taking under 20 minutes is (4+12)/40=16/40=0.4, or 40%.

Grouped data have already lost detail

If 16 observations lie in the interval 20–30, the histogram does not tell us their exact individual values. They could cluster near 20, near 30 or spread evenly.

Any grouped mean based on class midpoints is therefore an estimate, not the exact mean of the original raw data unless the actual values happen to align with those midpoints appropriately.

Worked Example 2 | Estimate a grouped mean

Using the travel-time data above, take midpoints 5, 15, 25 and 35.

Estimated total = 4(5)+12(15)+16(25)+8(35)
=20+180+400+280
=880.

Estimated mean = 880/40 = 22 minutes.

The word “estimated” matters because each class has been represented by its midpoint.

Dot plots preserve distribution shape

For a small data set, a dot plot shows repeated values, gaps, clusters and possible unusual observations directly. Two groups can have the same mean but visibly different spread.

For example, A={8,9,10,11,12} and B={2,6,10,14,18} both have mean 10. But B is much more spread out. A diagram that shows only the mean would hide this difference.

Stem-and-leaf diagrams preserve actual values

A stem-and-leaf diagram orders data while retaining individual observations. A key such as 4|7=47 tells the reader how to reconstruct values.

Because the raw values remain visible, the median, mode, range and quartiles can often be found directly without reconstructing an external list.

Worked Example 3 | Read a stem-and-leaf diagram

2 | 1 4 8
3 | 0 2 2 7
4 | 1 5
Key: 3 | 2 = 32

The data are 21,24,28,30,32,32,37,41,45.

There are 9 values, so the median is the 5th value: 32. The mode is also 32. The range is 45−21=24.

Cumulative frequency answers threshold questions

Cumulative frequency is a running total. It answers questions such as “how many observations are at or below this boundary?” and can be used to estimate median, quartiles and percentiles from grouped data.

The final cumulative frequency must equal the total sample size. If it does not, the running totals are wrong.

Worked Example 4 | Build cumulative totals

For frequencies 4,12,16,8 across successive intervals, the cumulative frequencies are:

  • 4;
  • 4+12=16;
  • 16+16=32;
  • 32+8=40.

The graph would plot these running totals against the upper class boundaries. The final point reaches total frequency 40.

Box plots compress five-number structure

A basic box plot summarises minimum, lower quartile, median, upper quartile and maximum. It is especially useful for comparing centre and spread between groups.

The box itself spans the interquartile range, which contains the middle half of the data under the quartile convention being used.

But a box plot does not show every individual data value. Two very different distributions can share the same five-number summary.

Worked Example 5 | Compare two box-plot summaries

Group A has median 68 and IQR 12. Group B has median 72 and IQR 24.

A supported comparison is:

Group B has the higher median, but its middle 50% is more spread out.

It would be unjustified to say every member of Group B scored higher than every member of Group A. The summaries do not support that claim.

Truncated axes can exaggerate differences

Suppose two values are 96 and 100. On an axis from 0 to 100, the bars look similar. On an axis beginning at 94, one bar can appear several times taller than the other even though the numerical difference is only 4 units.

A truncated axis is not automatically invalid. Sometimes it helps show small changes. But the reader should inspect the scale before accepting the visual impression.

Worked Example 6 | Quantify a dramatic-looking change

A chart shows a value rising from 200 to 210, but the vertical axis begins at 195. What is the actual percentage increase?

Percentage increase = (210−200)/200 ×100% = 5%.

The graphic may look dramatic because the visible axis interval is narrow, but the numerical change relative to the starting value is 5%.

Unequal visual area can distort category comparisons

If icons or pictures are enlarged in both height and width to represent a doubled value, their visual area can grow by a factor of four. A reader may perceive a much larger increase than the data justify.

When comparing pictorial displays, ask whether one dimension or total area is carrying the numerical encoding.

Pie charts require one common whole

A pie chart represents proportions of a total. Sector angle is:

sector angle = category frequency / total frequency × 360°.

Comparing sector sizes across two pie charts with different totals can be misleading if the actual counts are not considered. A 50% sector of 100 people represents 50 people; a 40% sector of 1000 represents 400 people.

Worked Example 7 | Percentage versus count

School A has 200 students and 60% join a club. School B has 500 students and 40% join the same type of club. Which has more club members?

School A: 0.60×200=120.
School B: 0.40×500=200.

School A has the higher percentage, but School B has the larger count. The correct comparison depends on the question.

Averages can hide distribution differences

Mean, median and mode answer different questions. A mean uses every value but can be pulled by extreme observations. A median describes the middle position and is more resistant to extreme values. Mode identifies most frequent values.

A display that reports only one average may omit important spread information. Whenever groups are compared, ask for both centre and variation where available.

Worked Example 8 | Same mean, different experience

Two delivery teams have times in minutes:

A: 28,29,30,31,32.
B: 10,20,30,40,50.

Both means are 30 minutes. But Team A is tightly clustered around 30, while Team B varies widely. Reporting only the common mean would conceal that operational difference.

Sample size matters

A percentage can appear impressive while being based on a very small sample. “80% preferred option A” means something different when 4 of 5 people were surveyed than when 800 of 1000 were surveyed.

The percentage alone does not reveal sample size, selection method or whether the sample represents the population of interest.

Correlation in a graph does not by itself establish cause

If two quantities rise together, the diagram can show association. It does not automatically prove that one caused the other. Other variables, reverse direction or coincidence may be relevant.

Good statistical interpretation distinguishes what is observed from what is inferred.

Worked Example 9 | State only what the display supports

A scatter-style display shows that students who report more study hours tend to have higher scores. A careful statement is: the data show a positive association between reported study hours and score in this sample.

The graph alone does not prove that increasing study hours by a particular amount will cause a particular score increase for every student.

Class intervals affect what the histogram reveals

The same raw data can look smoother or more irregular depending on how values are grouped. Very wide intervals hide local structure. Very narrow intervals can make random variation appear important.

When comparing histograms, check whether the class boundaries are actually the same. A visual comparison is weaker when one graph uses 5-unit intervals and another uses 20-unit intervals.

Worked Example 10 | Read a threshold from grouped frequencies

Using frequencies 4,12,16,8 for intervals 0–10, 10–20, 20–30, 30–40, how many observations are at least 20?

Use the last two complete classes:

16+8=24 observations.

Because 20 is exactly the lower boundary of the 20–30 class under the stated grouping, those observations are included.

A data-claim audit

  • Population: who or what is the claim about?
  • Sample: where did the data come from?
  • Measure: count, percentage, mean, median or spread?
  • Display: does the chosen graph fit the data type?
  • Scale: are axes truncated, uneven or visually exaggerated?
  • Grouping: have class intervals hidden important detail?
  • Comparison: are denominators, units and sample sizes comparable?
  • Inference: does the claim go beyond what the data show?

Common failure modes

ErrorCauseRepair
Calls a bar chart a histogramCategory and continuous data not distinguishedInspect variable type and whether intervals touch
Reads grouped midpoint as exact observationGrouping loss ignoredLabel grouped mean as estimated
Compares percentages without totalsDenominator ignoredConvert to counts when the question asks “how many”
Accepts dramatic bar-height impressionAxis start not inspectedRead numerical scale before visual comparison
Says higher median means every value is higherSummary statistic overextendedRestrict claim to centre information
Claims causation from associationInference exceeds evidenceUse association language unless causal evidence exists

Independent practice

  1. A histogram has equal-width classes with frequencies 5,9,14,12. Find the total frequency and modal class position by frequency.
  2. For intervals 0–10,10–20,20–30 with frequencies 6,10,4, estimate the mean using midpoints.
  3. A value rises from 80 to 84. Find the percentage increase.
  4. Group A has median 54 and IQR 8. Group B has median 50 and IQR 20. Write one supported comparison of centre and spread.
  5. Company X reports 70% satisfaction from 20 respondents. Company Y reports 65% satisfaction from 400 respondents. Find the number of satisfied respondents in each survey.
  6. A stem-and-leaf display represents 12,14,17,21,21,23,29. Find median, mode and range.

Explained answers

1. Total=5+9+14+12=40. The class with frequency 14 is the modal class.

2. Midpoints 5,15,25. Estimated total=6(5)+10(15)+4(25)=280. Divide by 20: estimated mean=14.

3. Increase=4. Percentage increase=4/80×100%=5%.

4. Group A has the higher median and the smaller IQR, so its middle half is less spread out.

5. X: 0.70×20=14. Y: 0.65×400=260.

6. Median=21, mode=21, range=29−12=17.

Teaching sequence: read before calculating

Begin with unlabeled chart types and ask learners what kind of variable each could represent. Then introduce scale audits: where does the axis start, what do intervals mean, and what information has been grouped away?

Next compare two groups using both centre and spread. Finish with short claims attached to diagrams and ask students to classify each claim as supported, unsupported or requiring more information.

Connect this guide to Cumulative Frequency, Box Plots and Standard Deviation, Statistics, Probability and Real-World Problems Under Exam Conditions, and Error Analysis, Corrections and Full-Paper Recovery.

Final thought

A statistical diagram is an argument made visible. Good mathematical reading checks the data type, scale, denominator, grouping and summary before accepting the visual story.

Do not ask only what the graph shows. Ask what choices made it look that way.

Return to the Secondary Mathematics Hub.