Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

Chapter 3: Statistical Data Analysis | Secondary 4 Mathematics Walkthrough | SEC G3 K310

Statistics is not only about obtaining a number. It is about deciding what a collection of numbers can reasonably tell us.

The supplied older Secondary 4 E-Mathematics textbook organises this chapter around cumulative frequency, median and quartiles, percentiles, range and interquartile range, box-and-whisker plots and standard deviation. That route remains strongly relevant to the current Singapore-Cambridge SEC G3 Mathematics syllabus, subject code K310. The current syllabus keeps these ideas and places them inside a wider expectation: students must choose, analyse and interpret statistical representations, recognise misleading presentations and compare data sets using appropriate measures.

This is an original eduKate learning walkthrough. The older textbook supplies a chapter map, not copied prose, worked examples or exercises. Current scope is determined by the 2027 SEC G3 K310 syllabus. Return to the Secondary 4 Mathematics Chapter-by-Chapter Walkthrough for the complete route.


SEC Check: What Is Still Tested?

Under K310, the statistical route includes analysis and interpretation of statistical representations, measures of central tendency, quartiles and percentiles, measures of spread, cumulative frequency diagrams, box-and-whisker plots and standard deviation. Students may also need to compare two sets of data using the mean and standard deviation and to judge whether a statistical representation is appropriate or potentially misleading.

  • tables and common statistical diagrams;
  • histograms, dot diagrams and stem-and-leaf diagrams;
  • cumulative frequency diagrams;
  • box-and-whisker plots;
  • mean, mode and median;
  • mean for grouped data;
  • quartiles and percentiles;
  • range and interquartile range;
  • standard deviation for grouped and ungrouped data;
  • comparison of data sets using centre and spread; and
  • purposes, strengths, weaknesses and possible misinterpretation of statistical displays.

The old chapter remains useful, but the current learning target is wider than “read a curve and calculate a standard deviation”. A student must connect representation, calculation and interpretation.

1. Begin With the Question the Data Is Supposed to Answer

Before calculating a mean or drawing a graph, ask what the data represents. A set of travel times, examination marks, temperatures and household expenses may all contain the same numerical values but require different interpretations. Statistics is meaningful only when the variable, units, population or sample and purpose are clear.

A good first reading asks:

  • What does each observation measure?
  • What are the units?
  • Is the data discrete or continuous?
  • Is the data grouped?
  • What comparison or conclusion is the question asking for?
  • Which representation preserves the information needed for that conclusion?

This prevents a common Secondary 4 mistake: performing a correct statistical procedure on the wrong quantity.

2. Cumulative Frequency Is About “Up to This Point”

A cumulative frequency total records how many observations are less than or equal to a chosen boundary. Instead of asking how many values lie inside one class interval, cumulative frequency asks how many have accumulated up to that boundary.

Suppose 40 students complete a task. If the cumulative frequency at 20 minutes is 27, then 27 students took no more than 20 minutes. The remaining 13 took longer than 20 minutes. The cumulative frequency is therefore a running total, not the frequency of the last class alone.

A cumulative frequency diagram lets the student estimate median, quartiles and percentiles visually. The vertical axis represents accumulated frequency. The horizontal axis represents the measured variable. The curve should move upward because a cumulative total cannot decrease as more of the distribution is included.

3. Median, Quartiles and Percentiles Locate Positions

The median divides an ordered data set into two halves. The lower quartile and upper quartile help divide it into quarters. Percentiles extend the same idea: a percentile identifies a position below which a stated percentage of the data lies.

For a cumulative frequency graph with 80 observations, the median is read near cumulative frequency 40. The lower quartile is read near 20 and the upper quartile near 60. A 90th percentile estimate is read near cumulative frequency 72. The precise graph-reading method depends on the scale, but the positional logic is stable.

Do not confuse “the 75th percentile” with “75% of the maximum value”. Percentiles describe position in the ordered data, not a percentage of the numerical scale.

4. Range and Interquartile Range Measure Spread Differently

The range compares the largest and smallest observations. It is simple and useful, but it is controlled entirely by the extremes. The interquartile range, usually written IQR, measures the width of the middle half of the data:

IQR = upper quartile − lower quartile.

Because the IQR concentrates on the middle 50%, it is less affected by very high or very low values than the full range. This difference matters when comparing data sets with outliers or unusually extreme observations.

Suppose Set A ranges from 10 to 90 but most values lie between 45 and 55, while Set B ranges from 30 to 70 and most values lie between 35 and 65. Set A has the larger full range, but Set B can still have the larger IQR. “More spread out” is therefore incomplete unless the measure of spread is named.

5. Box-and-Whisker Plots Compress a Distribution Into Five Landmarks

A box-and-whisker plot typically displays the minimum, lower quartile, median, upper quartile and maximum. The box shows the middle half of the data. The median marks the central position. The whiskers extend towards the extremes.

Box plots are especially useful for comparison because two distributions can be placed on the same scale. A student can compare:

  • median position;
  • interquartile range;
  • overall range;
  • relative symmetry or skew suggested by the spacing; and
  • whether one distribution appears consistently higher or more variable than another.

A good comparison uses both centre and spread. Saying “Class A did better because its median is higher” may be reasonable for typical performance, but it does not describe consistency. Saying “Class A is more consistent because its IQR is smaller” addresses spread but not level. A complete comparison identifies the measure used and what it implies.

6. Standard Deviation Measures Spread Around the Mean

Standard deviation measures how widely values tend to be distributed around the mean. A smaller standard deviation indicates that values are more tightly clustered around the mean. A larger standard deviation indicates greater spread.

For SEC work, the key is not merely obtaining the calculator value. The student must know what that value means. If two classes have similar means but one has a much smaller standard deviation, the class with the smaller standard deviation has more tightly clustered results around its mean.

K310 includes standard deviation for grouped and ungrouped data. Grouped data requires care because each class interval represents a range of possible values. When a grouped-data calculation uses class midpoints, the result is an estimate based on that representation. The working should preserve the distinction between exact raw observations and grouped approximations.

7. Compare Centre and Spread Together

Imagine two training groups:

GroupMean scoreStandard deviation
A724
B7511

Group B has the higher mean, so its average performance is higher. Group A has the smaller standard deviation, so its scores are more tightly clustered around its mean. Neither statement alone describes the whole distribution. A strong SEC response states both and connects each statistic to the interpretation it supports.

This is one reason statistical questions are not merely calculator questions. The marks often depend on reading what the measures mean in context.

8. Histograms and Other Representations Have Different Jobs

A table is good for exact values. A bar chart makes category comparisons visible. A line graph can show change over an ordered variable such as time. A histogram represents grouped continuous data through adjacent bars. A cumulative frequency graph supports positional estimates. A box plot gives a compact summary for comparison.

No representation is automatically “best”. The useful representation is the one that makes the required structure visible without distorting the data. K310 therefore expects students to understand purposes, uses, advantages and disadvantages of statistical representations, not simply recognise their names.

9. Misleading Data Displays: Ask What the Picture Is Making You Believe

A graph may use correct data and still create a misleading impression. A truncated vertical axis can exaggerate a small difference. Unequal intervals can distort apparent change. Three-dimensional decoration can make one bar appear disproportionately large. A graph that omits relevant categories or uses an unclear denominator can encourage the wrong comparison.

When asked whether a representation is misleading, identify the mechanism. “The graph is misleading” is not enough. A better explanation might be: “The vertical axis begins at 80 rather than 0, so a 4-point difference occupies most of the plotted height and appears much larger than it is.”

Worked Example: Read a Distribution Before Reaching for the Calculator

Two delivery teams record the time, in minutes, needed for a repeated route. Team P has a mean of 42 minutes and standard deviation 3 minutes. Team Q has a mean of 39 minutes and standard deviation 9 minutes.

Typical speed: Team Q is faster on average because its mean time is lower.

Consistency: Team P is more consistent because its standard deviation is smaller.

Interpretation: If a manager values the lowest average journey time, Team Q may look preferable. If predictable timing is important, Team P has the stronger evidence. Statistics does not make the decision automatically; it provides evidence for the criterion being used.

Common Failure Modes

  • Reading cumulative frequency as ordinary frequency: a cumulative value is a running total.
  • Reading the wrong axis: quartiles are located through cumulative frequency before reading the measured value.
  • Treating a percentile as a percentage of the maximum: it is a position in an ordered distribution.
  • Comparing only means: centre without spread can hide instability.
  • Comparing only standard deviations: spread without centre does not describe performance level.
  • Using “more consistent” without naming the measure: show whether the conclusion comes from IQR, standard deviation or another stated measure.
  • Entering grouped data incorrectly into a calculator: the values and frequencies must match the intended representation.
  • Calling any unusual graph misleading without explaining why: identify the scale, interval, omission or visual distortion.

A First-Principles Statistics Routine

  1. Name the variable. What is being measured?
  2. Check units and data type. Is the representation appropriate?
  3. Identify the question. Is it asking about centre, spread, position, shape or comparison?
  4. Choose the statistic or representation. Do not calculate everything by habit.
  5. Execute accurately. Use the calculator only after the data structure is correct.
  6. Interpret in context. Attach the conclusion to the measured variable.
  7. Check plausibility. Does the answer fit the scale and visible distribution?

How to Revise This Chapter in 30 Minutes

  1. 5 minutes: distinguish mean, median, quartiles, percentiles, range, IQR and standard deviation by purpose.
  2. 5 minutes: read median and quartiles from one cumulative frequency graph.
  3. 5 minutes: compare two box plots using centre and spread.
  4. 5 minutes: compute or retrieve a standard deviation and explain its meaning in one sentence.
  5. 5 minutes: inspect a misleading graph and name the exact distortion.
  6. 5 minutes: answer one comparison question without using vague words such as “better” or “more spread” unless the statistic is named.

Checkpoint Questions

  • Can I explain the difference between ordinary frequency and cumulative frequency?
  • Can I locate a median, quartile or percentile from cumulative frequency?
  • Can I distinguish range from IQR?
  • Can I read and compare box plots?
  • Can I explain what a standard deviation says about spread around the mean?
  • Can I compare two data sets using both mean and standard deviation?
  • Can I explain why a statistical diagram is misleading rather than merely state that it is?
  • Can I connect my numerical conclusion back to the context and units?

Continue the Learning Route

Next chapter: Chapter 4 — Matrices