SECONDARY 4 MATHEMATICS CLASSROOM · CHAPTER 3 · STATISTICAL DATA ANALYSIS · SEC G3 K310
Statistical Data Analysis: Read the Distribution Before You Reach for the Calculator
In this classroom, you will not begin by calculating every statistic you know. You will begin by asking what the data represents, what question is being asked, and which measure or diagram reveals the answer.
Statistics turns collections of observations into evidence. A table can preserve exact values. A histogram can reveal shape. A cumulative frequency diagram can reveal positions. A box plot can compare centre and spread. A mean can summarise level. A standard deviation can describe how tightly values cluster around that mean. The important skill is choosing the right view for the job.
Classroom rule: name the variable, units and question before choosing the statistic.
This classroom follows the current Singapore-Cambridge SEC G3 Mathematics syllabus, K310, where Data Handling and Analysis is listed under S1. The syllabus includes collecting, classifying and tabulating data; interpreting tables, bar graphs, pictograms, line graphs, pie charts, dot diagrams, histograms with equal class intervals, stem-and-leaf diagrams, cumulative frequency diagrams and box-and-whisker plots; evaluating statistical representations; drawing simple inferences; explaining misleading diagrams; mean, mode and median; grouped mean; quartiles and percentiles; range, interquartile range and standard deviation; standard deviation for grouped and ungrouped data; and comparing two data sets using mean and standard deviation.
Reference: 2027 SEC G3 syllabuses | SEAB.
Featured Answer: What Is Statistical Data Analysis?
Statistical data analysis is the process of organising, representing, summarising and interpreting data so that a sensible conclusion can be drawn. A correct calculation is only one part of that job. You must also know what the calculation means and whether it is the right calculation for the question.
A data set can be described through four broad questions:
- Centre: where is the data typically located?
- Spread: how variable is the data?
- Position: where does a particular observation sit within the distribution?
- Shape and representation: how is the data distributed, and which diagram reveals that well?
The Simple Classroom Answer
Statistics asks: what does this collection of observations say, and how strongly can we justify that conclusion?
Your first job is to identify the variable and units. Your second is to choose the representation or statistic. Your third is to interpret the result in context. Your fourth is to check whether the conclusion says more than the evidence supports.
How to Use This Classroom
- Read each teacher instruction.
- Pause at every Your Turn prompt.
- Write your interpretation before opening the worked answer.
- If your answer is wrong, identify whether the failure came from the data, the representation, the statistic, the arithmetic or the conclusion.
- Repeat with a changed example.
- Move to mixed examination transfer only after you can explain why each statistic is being used.
1. Start by Naming the Variable
Teacher: Put these five numbers on the board: 12, 14, 15, 18, 21. Ask the student, “What do these numbers mean?”
The correct answer is: we do not know yet. They could be ages, journey times, test scores, temperatures or numbers of books. Statistical meaning depends on context.
Now say: “These are waiting times in minutes.” Suddenly the same numbers support statements about service speed and variation.
Data without a variable is a list. Statistics begins when the list means something.
Your Turn 1
A table contains 42, 38, 51, 47 and 44. Give two completely different variables these values could represent.
Possible answer
They could represent examination marks out of 60, or journey times in minutes. The numbers are unchanged, but the interpretation and units change completely.
2. Identify the Units Before Calculating
Units tell you what a numerical statistic means. A mean of 42 is incomplete. A mean journey time of 42 minutes is interpretable. A standard deviation of 3 also inherits the units of the measured variable.
If two data sets use different units, convert them before direct comparison. Comparing 2.5 minutes with 130 seconds without conversion invites error.
3. Classify Before You Tabulate
Raw data is often reorganised into categories or intervals. The categories must be clear, non-overlapping and appropriate to the variable.
For example, if heights are grouped into 150–154 cm, 155–159 cm and 160–164 cm, each observation should belong to exactly one class. Ambiguous or overlapping intervals corrupt later frequencies.
Teacher check: ask the student why “10–20” and “20–30” can be ambiguous if class boundaries are not defined. The value 20 appears to belong to both groups unless the convention is made precise.
4. A Frequency Table Is a Counting Structure
A frequency table records how often each value or class occurs. It compresses repeated observations without losing the count.
| Score | Frequency |
|---|---|
| 1 | 2 |
| 2 | 4 |
| 3 | 5 |
| 4 | 3 |
| 5 | 1 |
Total frequency = 2 + 4 + 5 + 3 + 1 = 15. Any later mean, median or percentage calculation should be consistent with this total.
5. Choose the Diagram From the Question
K310 includes many statistical representations because different diagrams reveal different features.
| Representation | Best classroom use |
|---|---|
| Table | Preserve exact organised values. |
| Bar graph | Compare separate categories. |
| Pictogram | Show category frequency visually. |
| Line graph | Show change across an ordered variable such as time. |
| Pie chart | Show parts of a whole. |
| Dot diagram | See individual values and clustering in smaller data sets. |
| Histogram with equal class intervals | Show the distribution of grouped continuous data. |
| Stem-and-leaf diagram | Show distribution while preserving individual values. |
| Cumulative frequency diagram | Read medians, quartiles and percentiles. |
| Box-and-whisker plot | Compare central position and spread compactly. |
The correct question is not “Which diagram do I remember?” It is “Which diagram makes the required feature easiest to see?”
6. Bar Graphs and Histograms Are Not the Same Thing
A bar graph usually represents separate categories, so gaps between bars can be meaningful. A histogram represents grouped continuous data, so the bars touch because adjacent class intervals cover a continuous scale.
For the K310 histogram content, equal class intervals are specified. This means direct bar-height comparison is straightforward because class widths are the same.
Teacher: Draw a bar chart of favourite subjects and a histogram of journey times. Ask the student why gaps make sense in one but not the other.
7. Stem-and-Leaf Diagrams Preserve Individual Values
A stem-and-leaf diagram groups values while retaining each observation. If the data are 21, 24, 27, 31, 33, 33 and 38, the tens digits can form stems and the units digits leaves.
The diagram reveals clustering and spread while allowing the original data to be reconstructed. Always include a key so the reader knows what a stem and leaf mean.
8. Dot Diagrams Make Repeated Values Visible
A dot diagram places one mark for each observation along a number scale. Repeated values stack vertically. This makes clusters, gaps and repeated values easy to see without compressing the data into a single statistic.
A student should learn to read a distribution before summarising it. A mean alone can hide a cluster, a gap or an outlier that is obvious in the dot diagram.
9. Mean, Median and Mode Are Different Views of Centre
The three common measures of central tendency answer different questions.
- Mean: the equal-share or arithmetic average, using every numerical observation.
- Median: the middle position after ordering the data.
- Mode: the most frequently occurring value or category.
Do not ask which measure is “best” without context. Ask which measure is most useful for the decision or description required.
10. Teacher Model 1: Mean From Raw Data
Five test scores are 12, 15, 18, 20 and 25.
Total = 12 + 15 + 18 + 20 + 25 = 90.
Number of scores = 5.
Mean = 90 ÷ 5 = 18.
The mean uses every score. Changing one score can change the mean even if the median remains fixed.
Your Turn 2
Find the mean of 8, 11, 13, 14 and 19.
Answer
Total = 65. Mean = 65 ÷ 5 = 13.
11. Mean From a Frequency Table
When values repeat, multiply each value by its frequency before adding.
| x | f | fx |
|---|---|---|
| 1 | 2 | 2 |
| 2 | 4 | 8 |
| 3 | 5 | 15 |
| 4 | 3 | 12 |
| 5 | 1 | 5 |
Σf = 15 and Σfx = 42.
Mean = 42/15 = 2.8.
The frequency column tells you how many times each value contributes.
12. Grouped Mean Uses Representative Class Values
K310 includes calculation of the mean for grouped data. When exact values inside each class are not known, use a representative value such as the class midpoint.
| Time, t (min) | Frequency | Midpoint | f × midpoint |
|---|---|---|---|
| 0–9 | 3 | 4.5 | 13.5 |
| 10–19 | 5 | 14.5 | 72.5 |
| 20–29 | 2 | 24.5 | 49 |
Total frequency = 10. Total estimated value = 135.
Estimated mean = 135 ÷ 10 = 13.5 minutes.
The word estimated matters because the midpoint stands in for unknown individual observations inside each interval.
13. Do Not Pretend Grouped Data Is Exact Raw Data
Suppose the class 10–19 minutes contains five observations. The midpoint method treats all five as 14.5 minutes for the grouped calculation. The actual values could be 10, 11, 13, 18 and 19. Grouping has compressed the information.
This does not make grouped analysis useless. It means the student should understand what has been approximated.
14. Median Requires Ordered Position
To find the median of raw data, order the observations first. The median is a positional statistic. If there is an odd number of observations, identify the middle one. If there is an even number, take the mean of the two middle observations where appropriate.
For 4, 7, 9, 11, 18, the median is 9.
For 4, 7, 9, 11, 18, 23, the middle pair is 9 and 11, so the median is 10.
15. Mode Is About Frequency, Not Size
The mode is the most frequently occurring value or category. It is not the largest observation.
In 2, 2, 3, 5, 5, 5, 9, the mode is 5 because it occurs most often.
For categorical data such as preferred transport mode, the mode can still be meaningful even when a numerical mean is not.
16. An Extreme Value Can Pull the Mean
Consider monthly incomes in thousands: 3, 3, 4, 4, 36.
Mean = 50/5 = 10. Median = 4.
The mean is much higher than most observations because the value 36 pulls it upward. In this data set, the median may better describe a typical central position.
The lesson is not “median is always better when there is an outlier”. The lesson is that choice of centre depends on the question and distribution.
17. Cumulative Frequency Means Running Total
A cumulative frequency records how many observations have accumulated up to a boundary.
| Time below | Frequency in class | Cumulative frequency |
|---|---|---|
| 10 | 4 | 4 |
| 20 | 7 | 11 |
| 30 | 5 | 16 |
| 40 | 4 | 20 |
The cumulative frequency 16 at 30 means 16 observations are below the stated boundary. It does not mean the third class itself contains 16 observations.
18. Build Cumulative Frequency by Accumulating
If ordinary frequencies are 3, 5, 8, 4, then cumulative frequencies are:
- 3
- 3 + 5 = 8
- 8 + 8 = 16
- 16 + 4 = 20
The final cumulative frequency must equal the total number of observations. This is an immediate check.
Your Turn 3
Ordinary frequencies are 6, 9, 7, 3. Write the cumulative frequencies.
Answer
6, 15, 22, 25. The final cumulative frequency 25 is the total number of observations.
19. A Cumulative Frequency Curve Should Never Fall
As the upper boundary increases, more observations can be included, never fewer. Therefore a cumulative frequency curve should move upward or stay level. A downward section signals a plotting or reading error.
The vertical axis is cumulative frequency. The horizontal axis is the measured variable.
20. Median on a Cumulative Frequency Diagram
If there are 80 observations, the median lies around the 40th observation. On the cumulative frequency axis, locate 40, move horizontally to the curve, then drop vertically to the measurement axis.
Position first on the cumulative-frequency axis; value second on the measurement axis.
This direction of reading prevents a common axis error.
21. Quartiles Divide Ordered Data Into Four Positional Regions
The lower quartile Q1 lies around the 25th percentile. The median Q2 lies around the 50th percentile. The upper quartile Q3 lies around the 75th percentile.
For 120 observations:
- Q1 is read near cumulative frequency 30;
- median near 60;
- Q3 near 90.
These positions describe how the ordered data is partitioned, not percentages of the maximum numerical value.
22. Percentiles Generalise the Same Positional Idea
The 90th percentile is the value below which about 90% of the observations lie. In a data set of 200 observations, locate cumulative frequency 180 and read the corresponding variable value.
Do not calculate 90% of the largest measurement. Percentiles refer to position in the ordered distribution.
23. Teacher Model 2: Read Positions From a Cumulative Frequency Graph
Suppose a cumulative frequency graph represents 60 journey times.
- Median: locate cumulative frequency 30.
- Lower quartile: locate 15.
- Upper quartile: locate 45.
- 80th percentile: locate 48.
After locating each position on the vertical axis, move to the curve and down to the journey-time axis. The resulting values are estimates read from the graph.
24. Range Measures the Full Width of the Data
Range = maximum − minimum.
For 5, 7, 10, 11, 18, the range is 18 − 5 = 13.
The range is easy to calculate but depends completely on two extreme values. One unusual observation can change it dramatically.
25. Interquartile Range Measures the Middle Half
IQR = Q3 − Q1.
The IQR measures the width of the middle 50% of the distribution. It ignores the lowest quarter and highest quarter when forming the width, so it is less controlled by extreme values than the full range.
If Q1 = 18 and Q3 = 31, then IQR = 13.
26. “More Consistent” Must Be Tied to a Measure
If two groups have IQRs 8 and 15, the group with IQR 8 has a narrower middle half. You can describe it as more consistent with respect to the IQR.
If the comparison uses standard deviation instead, use the standard deviation. Do not switch measures silently.
27. Box-and-Whisker Plots Compress Five Key Positions
A box-and-whisker plot typically shows:
- minimum;
- lower quartile Q1;
- median;
- upper quartile Q3;
- maximum.
The box spans Q1 to Q3, so its width is the IQR. The line inside the box marks the median. The whiskers extend towards the extremes.
28. Read the Box Plot as Centre and Spread
When comparing two box plots, first compare medians for central position. Then compare IQRs for the spread of the middle half. If useful, compare overall ranges as well.
A strong comparison might say:
Group B has the higher median, so its typical score is higher, while Group A has the smaller IQR, so the middle half of its scores is more tightly clustered.
This is stronger than “B is better” because it separates level from consistency.
29. Teacher Model 3: Compare Two Box Plots Numerically
Group A: median 68, Q1 = 60, Q3 = 74.
Group B: median 72, Q1 = 58, Q3 = 82.
Group A IQR = 74 − 60 = 14.
Group B IQR = 82 − 58 = 24.
Group B has the higher median. Group A has the smaller IQR and therefore a tighter middle 50%.
30. Standard Deviation Measures Spread Around the Mean
Standard deviation describes how widely observations are distributed around the mean. A smaller standard deviation means observations are more tightly clustered around the mean. A larger standard deviation means greater spread around the mean.
The standard deviation has the same units as the original variable. If the data is measured in seconds, the standard deviation is in seconds.
31. Standard Deviation Is Not “Average Distance” in a Casual Sense
At school level, you should understand standard deviation as a formal measure of spread around the mean rather than inventing an informal shortcut. Use the calculator or required method correctly, then interpret the resulting value.
The crucial interpretation is comparative: smaller standard deviation means tighter clustering around the mean, provided the data sets are being compared appropriately.
32. Teacher Model 4: Same Mean, Different Spread
Data Set A: 48, 49, 50, 51, 52.
Data Set B: 30, 40, 50, 60, 70.
Both means are 50. But B is far more spread out. Therefore B must have the larger standard deviation.
You can predict the comparison before using a calculator by inspecting the distances from the common mean.
33. Mean and Standard Deviation Must Be Interpreted Together
| Group | Mean | Standard deviation |
|---|---|---|
| A | 72 | 4 |
| B | 75 | 11 |
Group B has the higher average score. Group A has the smaller standard deviation, so its scores are more tightly clustered around its mean.
Neither fact should be erased by the other. Statistical comparison often requires two statements because centre and spread answer different questions.
34. Do Not Say “Better” Until You Define Better
If Group B has a higher mean but Group A is more consistent, which is “better”? The mathematics alone cannot answer without a criterion.
If the goal is highest average performance, the mean matters. If the goal is predictability, spread matters. A good conclusion ties the statistic to the decision criterion.
35. Grouped Standard Deviation Is Based on Grouped Representation
K310 includes standard deviation for grouped and ungrouped data. For grouped data, representative class values such as midpoints are used in the calculation where exact raw observations are unavailable.
That means the grouped calculation summarises the grouped model of the data. Keep the class midpoints and frequencies aligned correctly when entering the calculator.
36. Calculator Discipline: Enter the Data Structure, Not Just the Numbers
For ungrouped raw data, enter each value correctly. For a frequency table, ensure each value is matched with its frequency. For grouped data, use the intended representative class value and its frequency.
Before trusting the standard deviation, verify the calculator’s reported count or total frequency where possible. A beautiful decimal from a wrongly entered data set is still wrong.
37. Misleading Graphs: Ask What Visual Claim the Diagram Is Making
A graph can use true numbers and still create a misleading impression. K310 explicitly includes explaining why a statistical diagram leads to misinterpretation.
Inspect:
- axis starting points;
- unequal or unclear scales;
- missing labels or units;
- distorted pictogram sizes;
- inconsistent intervals;
- omitted categories;
- 3D decoration that exaggerates visual size;
- whether the denominator or total is hidden.
38. Teacher Model 5: Truncated Axis
Two values are 94 and 98. A bar graph begins its vertical axis at 90. The difference of 4 units may occupy most of the visible bar height, making the second value look several times larger.
A strong explanation is:
The vertical axis begins at 90 instead of showing the full scale from zero, so the 4-unit difference occupies a disproportionately large part of the displayed height and exaggerates the visual difference.
“The graph is misleading” without the mechanism is not enough.
39. Pictograms Can Distort Area
If one icon is doubled in height and doubled in width, its visible area becomes four times as large, not twice. A pictogram that scales both dimensions can make a twofold numerical increase look fourfold visually.
Ask what property of the icon is supposed to represent quantity: count, length or area.
40. A Pie Chart Is About Proportion of a Whole
A full circle represents the whole data set. Sector angle is proportional to category frequency.
If 30 out of 120 students choose a category, the fraction is 30/120 = 1/4, so the sector angle is 1/4 × 360° = 90°.
Do not compare raw sector angles across two pie charts unless the total sizes and intended comparison are understood. A larger sector in a smaller population can still represent fewer people.
41. Line Graphs Need an Ordered Horizontal Variable
A line graph is especially useful when the horizontal variable has a meaningful order, often time. Connecting points suggests movement or change across that order.
Do not connect unrelated categories with a line simply because the software allows it. Representation should match meaning.
42. Draw Simple Inferences, Not Unlimited Conclusions
K310 includes drawing simple inference from statistical diagrams. An inference should follow from the displayed evidence.
If a line graph shows higher sales every weekend in the observed period, you may state that weekend sales were higher during that period. You should not automatically claim that weekends cause higher sales or that the pattern must continue forever.
Evidence supports conclusions within limits.
43. Misconception Clinic: Cumulative Frequency Is Ordinary Frequency
If cumulative frequency at 30 is 42, that means 42 observations have accumulated up to that boundary. It does not mean the 20–30 class contains 42 observations.
To recover a class frequency from cumulative frequencies, subtract consecutive cumulative totals.
44. Misconception Clinic: Percentile Means Percentage of the Maximum
The 80th percentile is a positional value. About 80% of observations lie at or below that level. It is not 80% of the maximum measurement.
Correction routine: identify the total number of observations, calculate 80% of that count, then use the cumulative-frequency position.
45. Misconception Clinic: Smaller Mean Means More Consistent
The mean describes centre, not consistency. Consistency is a spread question. Use IQR or standard deviation depending on the information given.
A data set can have a low mean and huge spread, or a high mean and tiny spread.
46. Misconception Clinic: Smaller Range Always Means Smaller Standard Deviation
Range uses only the maximum and minimum. Standard deviation uses the distribution of all observations around the mean. Two data sets can have the same range but very different standard deviations.
Do not infer one spread measure directly from another without evidence.
47. Misconception Clinic: Same Mean Means Same Distribution
Data sets A = {48,49,50,51,52} and B = {30,40,50,60,70} both have mean 50. Their distributions are clearly different.
Centre is not the whole distribution. Always inspect spread and representation where relevant.
48. Misconception Clinic: Grouped Mean Is Exact
If exact values inside classes are unknown, a midpoint-based grouped mean is an estimate. Write or interpret it accordingly.
The midpoint stands in for all observations in that class during the calculation.
49. Misconception Clinic: “More Spread” Without Naming the Measure
Say whether the conclusion comes from range, IQR or standard deviation.
A precise statement is stronger:
Group B has the larger standard deviation, so its values are more widely spread around its mean.
50. Misconception Clinic: Calculator Output Is Automatically Correct
A calculator only processes the data you enter. Wrong frequencies, wrong midpoints or missing observations can still produce a neat-looking result.
Check total frequency, rough mean location and plausible spread before accepting the output.
50A. K310 Transfer Ladder: AO1 → AO2 → AO3
Statistical data analysis is one of the clearest places to train the full K310 assessment progression because the same data can first be summarised, then interpreted, then challenged.
| Assessment mode | Statistics task | What a strong response shows |
|---|---|---|
| AO1 | calculate or read mean, median, quartiles, IQR, standard deviation, cumulative frequency and statistical diagrams | correct data entry, arithmetic, graph reading, units and notation |
| AO2 | compare two real or simulated data sets and choose which measures and representations answer the actual question | relevant statistic selection, centre/spread separation, context-aware comparison and sensible inference |
| AO3 | justify why one conclusion is stronger, explain why a graph is misleading, or state why a claim goes beyond the evidence | a complete argument that names the statistical mechanism and its effect |
Teacher progression: one AO1 calculation → one AO2 comparison → one AO3 critique. Then present a mixed data question without naming the statistic so the learner must decide what evidence is sufficient.
AO2 Transfer Example: Commute Reliability
Route A has mean journey time 38 minutes and standard deviation 3 minutes. Route B has mean journey time 35 minutes and standard deviation 9 minutes. A student wants the fastest route on average, while another values predictable arrival time. Explain how the recommendation changes.
Worked transfer
Route B has the lower mean, so it is faster on average. Route A has the smaller standard deviation, so its journey times are more tightly clustered around its mean and it is more consistent. The recommendation therefore depends on whether average speed or predictability is the decision criterion.
AO3 Reasoning Example: Misleading Diagram
A graph compares values 96 and 100, but the vertical axis begins at 95. A student says the second value is “about five times as large” because its bar looks five times taller. Explain the error.
Reasoning answer
The axis is truncated close to the data values, so the displayed bar heights represent only the differences above 95 rather than the full values. The visual height ratio therefore does not equal the numerical value ratio. The actual values 96 and 100 differ by only 4, so the graph exaggerates the apparent relative difference.
AO3 Reasoning Example: Evidence Limits
A sample of 40 students shows that students who slept longer had higher test scores. Explain why this does not by itself prove that extra sleep caused the higher scores.
Reasoning answer
The data shows an association within the observed sample, but other variables may differ between the students, and the study design may not control those factors. Statistical association alone does not establish a causal relationship. A safer conclusion is that longer sleep was associated with higher scores in this sample.
51. Guided Practice Set A: Centre
Data: 4, 6, 6, 8, 11.
- Find the mean.
- Find the median.
- Find the mode.
- Find the range.
Solutions
Mean = 35/5 = 7. Median = 6. Mode = 6. Range = 11 − 4 = 7.
52. Guided Practice Set B: Frequency Table Mean
| x | f |
|---|---|
| 2 | 3 |
| 3 | 4 |
| 4 | 2 |
| 5 | 1 |
Find the mean.
Worked solution
Σf = 10. Σfx = 2(3) + 3(4) + 4(2) + 5(1) = 31. Mean = 31/10 = 3.1.
53. Guided Practice Set C: Grouped Mean
| Mass (kg) | Frequency |
|---|---|
| 40–49 | 2 |
| 50–59 | 5 |
| 60–69 | 3 |
Use midpoints 44.5, 54.5 and 64.5 to estimate the mean.
Worked solution
Weighted total = 44.5(2) + 54.5(5) + 64.5(3) = 89 + 272.5 + 193.5 = 555. Total frequency = 10. Estimated mean = 55.5 kg.
54. Guided Practice Set D: Cumulative Frequency
Ordinary class frequencies are 5, 8, 12, 9 and 6.
- Write the cumulative frequencies.
- State the total number of observations.
- Which cumulative position would be used for the median?
Solutions
Cumulative frequencies: 5, 13, 25, 34, 40. Total = 40. The median is read around cumulative frequency 20.
55. Guided Practice Set E: Quartiles and IQR
A box plot has Q1 = 24, median = 31 and Q3 = 39.
- Find the IQR.
- What percentage of the data lies approximately between Q1 and Q3?
- What does the median 31 mean positionally?
Solutions
IQR = 39 − 24 = 15. About 50% of the observations lie between Q1 and Q3. The median is the central positional value, with about half the observations at or below it and half at or above it.
56. Guided Practice Set F: Compare Two Data Sets
| Team | Mean time | Standard deviation |
|---|---|---|
| P | 42 min | 3 min |
| Q | 39 min | 9 min |
- Which team is faster on average?
- Which team is more consistent?
- Write a two-sentence comparison.
Worked answer
Q is faster on average because its mean time is lower. P is more consistent because its standard deviation is smaller, so its times are more tightly clustered around its mean.
57. Guided Practice Set G: Misleading Graph
A bar chart compares scores 88 and 92 but the vertical axis begins at 86.
Explain why the graph may mislead.
Worked answer
The axis is truncated close to the data values, so the 4-point difference occupies a large proportion of the displayed height and visually exaggerates the difference between the scores.
58. Challenge Practice: Same Mean, Different Interpretation
Class A and Class B both have mean 70. Class A has standard deviation 5 and Class B has standard deviation 14.
- What can you say about average performance?
- What can you say about consistency?
- Can you conclude which class has the higher median?
Worked answer
The classes have the same mean, so their average scores are equal. Class A is more tightly clustered around its mean because its standard deviation is smaller. The information given is not enough to determine which class has the higher median.
59. Challenge Practice: Recover a Class Frequency
A cumulative frequency table has totals 7, 19, 31 and 40 at successive class boundaries.
Find the ordinary frequency in each class.
Solution
First class = 7. Second = 19 − 7 = 12. Third = 31 − 19 = 12. Fourth = 40 − 31 = 9. Frequencies: 7, 12, 12, 9.
60. Challenge Practice: Compare Two Box Plots
Box Plot A: median 62, Q1 55, Q3 69. Box Plot B: median 66, Q1 50, Q3 78.
Worked comparison
B has the higher median, so its central score is higher. A has IQR 14 while B has IQR 28, so A’s middle 50% is much more tightly clustered.
61. Challenge Practice: Percentage to Pie-Chart Angle
Thirty-five percent of students choose option A. Find the corresponding pie-chart angle.
Solution
0.35 × 360° = 126°.
62. Examination Method: Write What Each Statistic Means
After finding a mean or standard deviation, attach a phrase:
- “The mean journey time is 42 minutes.”
- “The smaller standard deviation indicates more tightly clustered journey times around the mean.”
- “The median score is higher, indicating a higher central position.”
This prevents a correct number from being attached to the wrong interpretation.
63. Examination Method: Compare Centre and Spread Separately
Use a two-part comparison:
- compare mean or median for central level;
- compare standard deviation, IQR or range for spread.
Do not compress both ideas into “A is better”.
64. Examination Method: Use the Correct Axis First
For cumulative frequency, start with the required position on the cumulative-frequency axis. Then move to the curve and read the measurement.
For a line graph or other graph, first read the axis labels and units. Never assume the horizontal axis is time or the vertical axis is frequency.
65. Examination Method: Show the Grouped Mean Structure
If calculating a grouped mean by hand, show class midpoints and frequency products clearly. This makes the estimate auditable and reduces midpoint-entry errors.
Use:
Estimated mean = Σ(f × representative value) ÷ Σf.
66. Examination Method: Name the Misleading Mechanism
If a graph is misleading, state exactly why:
- truncated axis;
- unequal scale;
- distorted icon size;
- missing category;
- unclear denominator;
- inappropriate representation.
Then explain the visual effect on the reader.
67. Examination Method: Keep Inference Inside the Evidence
Use phrases such as “the data suggests”, “in this sample” or “during the observed period” when the evidence does not justify a universal conclusion.
Statistics summarises evidence. It does not automatically prove causation or certainty about future cases.
68. Oral Classroom Check
- What is the difference between a bar graph and a histogram?
- When is a stem-and-leaf diagram useful?
- Why can the mean be affected by an extreme value?
- What does cumulative frequency mean?
- How do you locate the median on a cumulative frequency graph?
- What is the difference between range and IQR?
- What does standard deviation describe?
- Why must grouped mean sometimes be described as estimated?
- How do you compare two data sets using mean and standard deviation?
- What makes a graph misleading?
The student should answer in complete mathematical sentences and use examples. If an answer is only a memorised formula, ask what the formula means in context.
69. Exit Ticket
Two groups complete the same task.
| Group | Mean time | Standard deviation |
|---|---|---|
| A | 36 min | 4 min |
| B | 33 min | 10 min |
- Which group is faster on average?
- Which group is more consistent?
- If predictable timing is more important than lowest average time, which statistic supports Group A?
- Can you conclude that every member of Group B was faster than every member of Group A?
Exit-ticket solution
B is faster on average because its mean time is lower. A is more consistent because its standard deviation is smaller. The standard deviation supports A for predictability. No: the mean and standard deviation do not prove that every B value is lower than every A value.
70. Homework: Retrieval, Variation and Transfer
Layer 1 — Retrieval
- Define mean, median, mode, range, IQR and standard deviation.
- Write the five landmarks of a box plot.
- Explain cumulative frequency in one sentence.
- Name four reasons a statistical diagram may mislead.
Layer 2 — Variation
- One raw-data mean/median/mode question.
- One frequency-table mean question.
- One grouped-mean question.
- One cumulative-frequency quartile question.
- One box-plot comparison.
- One mean-and-standard-deviation comparison.
- One misleading-graph explanation.
Layer 3 — Transfer
Collect a small real data set such as commute times, study durations or daily temperatures. Choose one representation, calculate one measure of centre and one measure of spread, then write a short conclusion. Finally, explain one limitation of the data or conclusion.
71. The Full Chapter Routine
For any statistics question, use:
Variable → units → representation → question → statistic → calculation → interpretation → limitation/check.
For cumulative frequency, use:
Total → positional frequency → curve → measurement value → interpret.
For comparison, use:
Centre first → spread second → context conclusion.
For misleading graphs, use:
Name the mechanism → explain the visual distortion → state the safer interpretation.
72. Why This Chapter Matters Beyond Statistics
Statistical reasoning trains you to separate evidence from impression. It asks whether the representation is fair, whether the summary hides important variation, whether two groups differ in centre or spread, and whether a conclusion is justified by the observed data.
These habits matter in science, economics, medicine, engineering, business, public policy, sport, education and everyday decision-making. Numbers become useful when their meaning, uncertainty and limitations remain visible.
73. Connect Back to Probability
Chapter 2 models possible outcomes before they occur. Chapter 3 analyses observations after they have occurred. Both chapters deal with uncertainty, but they work from different directions.
If chance-event structure is weak, return to the Chapter 2 Combined Probability Classroom. If data representation and interpretation are weak, stay here until centre, spread and position become separate ideas.
74. Ready for Chapter 4?
You are ready to move on when you can do all of the following without prompts:
- name the variable and units;
- choose a suitable statistical representation;
- distinguish bar graphs from histograms;
- calculate and interpret mean, median and mode;
- calculate a grouped mean using representative values;
- construct cumulative frequency totals;
- read median, quartiles and percentiles from cumulative frequency;
- calculate and interpret range and IQR;
- read and compare box plots;
- interpret standard deviation as spread around the mean;
- compare two data sets using mean and standard deviation;
- identify why a statistical diagram is misleading; and
- write a conclusion that stays inside the evidence.
If one item is weak, return to that section and complete a changed example. If all are stable, continue to Matrices, where information is organised into rectangular structures and operations depend on row-column meaning.
Continue the Secondary 4 Mathematics Classroom
- Secondary 4 Mathematics Chapter-by-Chapter Walkthrough | SEC G3 K310
- Secondary 4 Mathematics Learning Guide | Cumulative Frequency, Box Plots and Standard Deviation
- Secondary 4 Mathematics Learning Guide | Mean, Median, Mode, Range and Comparing Data Sets
- Secondary 4 Mathematics Learning Guide | Histograms, Statistical Diagrams and Misleading Data
- Previous: Chapter 2 | Probability of Combined Events
- Next: Chapter 4 | Matrices
