Statistics begins before the graph is drawn. The quality of a conclusion depends on what was measured, who was measured, how values were classified, how observations were tabulated and whether the chosen representation suits the data.
This fortieth Secondary 4 Mathematics Learning Guide develops data collection, classification, tabulation and representation choice as one statistical reasoning system. It belongs to the Secondary Mathematics Hub and S1–S4 Capability Map.
It complements Histograms, Statistical Diagrams and Misleading Data, Mean, Median, Mode, Range and Comparing Data Sets and Cumulative Frequency, Box Plots and Standard Deviation.
Start by identifying the variable
A statistical variable records a characteristic that can vary from one observation to another.
| Variable type | Examples |
|---|---|
| Categorical | transport mode, house type, favourite subject |
| Discrete numerical | number of siblings, books borrowed, goals scored |
| Continuous numerical | height, mass, time, temperature |
The variable type influences which table, summary and diagram are meaningful.
Worked Example 1 | Classify variables
Classify each variable:
- Number of messages received in one hour → discrete numerical.
- Time taken to complete a task → continuous numerical.
- Preferred study location → categorical.
Population and sample answer different questions
The population is the full group about which we want information. A sample is the subset actually measured.
A sample can be easier and cheaper to collect, but its conclusions are only useful if the sample reasonably represents the population relevant to the question.
Worked Example 2 | Identify population and sample
A school wants to estimate the average travel time of all 1200 students and surveys 150 students.
- Population: all 1200 students.
- Sample: the 150 surveyed students.
Sampling method affects credibility
A convenience sample may over-represent people who are easiest to reach. A voluntary-response sample may over-represent people with strong opinions. A carefully selected sample aims to reduce systematic bias.
Mathematics questions may not require a full survey-design theory, but they often expect the learner to recognise when a method could produce unrepresentative data.
Worked Example 3 | Detect sampling bias
A canteen asks only students currently buying vegetarian food whether the school should increase vegetarian options.
This sample is likely biased toward students who already choose vegetarian food. It does not represent all canteen users well enough for a school-wide conclusion.
Raw data become useful through classification
Raw observations may be unsorted and difficult to interpret. Classification groups values or categories so frequency can be seen.
For categorical data, classification may simply count each category. For continuous numerical data, class intervals may be needed.
Worked Example 4 | Build a frequency table
The raw values are 2,3,3,5,2,4,3,5,5,3.
| Value | Frequency |
|---|---|
| 2 | 2 |
| 3 | 4 |
| 4 | 1 |
| 5 | 3 |
Total frequency=10, matching the number of raw observations. This is a useful verification check.
Class intervals must be unambiguous
Intervals such as 0≤t<10, 10≤t<20 and 20≤t<30 avoid overlap. Each observation belongs to exactly one class.
If one class were written 0–10 and the next 10–20 without a convention, a value exactly equal to 10 could be ambiguous.
Worked Example 5 | Place boundary values correctly
Classes are 0≤x<5, 5≤x<10 and 10≤x<15. Where does x=10 belong?
10≤x<15.
The lower boundary is included; the upper boundary of the previous class is excluded.
Choose a bar chart for separate categories
Bar charts are useful for comparing distinct categories. The gaps help signal that one category does not flow continuously into the next.
Category order may be natural, alphabetical or chosen for comparison, depending on context.
Choose a line graph for ordered change
A line graph is useful when the horizontal variable has an order, commonly time. Connecting the plotted points highlights progression from one observation to the next.
It is usually inappropriate to connect unrelated categories merely because they can be listed from left to right.
Worked Example 6 | Choose between bar and line graph
Which display is more suitable?
- Daily temperature over seven consecutive days → line graph.
- Number of students choosing English, Mathematics, Science and Art → bar chart.
Pie charts show parts of one whole
A pie chart is appropriate when categories partition a total and proportion is central to the question.
Sector angle = category frequency / total frequency × 360°.
Worked Example 7 | Sector angle
In a survey of 120 students, 30 choose option A. Find the sector angle.
30/120×360°=90°.
Histograms represent grouped continuous data
Histogram bars touch because class intervals lie on a continuous numerical scale. This differs from a bar chart, where categories are separate.
For the equal-class-width histograms used in the current core syllabus, frequency can be represented directly by bar height when widths are equal.
Dot plots preserve small-data detail
For a small numerical data set, a dot plot can show every observation while making clusters, gaps and repeated values visible.
It is often more informative than compressing a small data set immediately into broad class intervals.
Stem-and-leaf preserves individual values and order
A stem-and-leaf diagram keeps actual data values while sorting them. It can therefore support direct calculation of median, mode, range and quartiles.
Worked Example 8 | Choose a detailed display
A teacher has 18 test scores and wants a display that preserves each exact score while showing distribution shape. A stem-and-leaf diagram or dot plot is more suitable than a pie chart.
Cumulative frequency answers threshold and percentile questions
Cumulative frequency is useful when the question asks how many values lie below a threshold, or when estimating median, quartiles or percentiles from grouped data.
The final cumulative total should equal the sample size.
Box plots support compact group comparison
A box plot shows a five-number summary and allows quick comparison of median and spread between groups.
Its strength is compact comparison; its weakness is that individual observations and detailed distribution shape are hidden.
Worked Example 9 | Choose a display for two groups
Two classes each have 40 scores. The goal is to compare median and spread compactly. A pair of box plots is suitable.
If the goal instead were to inspect every individual score, a box plot would hide too much detail.
Representation choice depends on the question, not just the data
The same data set may support more than one valid display. A line graph may emphasise trend over time; a table may make exact values easiest to retrieve; a box plot may make comparison of spread clearer.
Ask what the reader needs to see.
Misleading scales can distort perception
An axis that begins close to the smallest data value can make a small difference look visually dramatic. Unequal intervals can also distort comparison if not clearly indicated.
A reader should inspect the numerical scale before trusting the visual impression.
Worked Example 10 | Truncated-axis claim
A bar rises from 96 to 100, and the graph axis begins at 95. What is the numerical percentage increase?
(100−96)/96×100%≈4.17%.
The visual bar-height difference may look much larger because the axis is truncated.
Question wording can bias responses
A survey question such as “Do you agree that the improved new system should be kept?” contains persuasive wording. A more neutral question would avoid telling respondents that the system is improved.
Data quality begins with measurement design, not only with later calculation.
Missing data and non-response can matter
If many selected participants do not respond, the final respondents may differ systematically from those who did not. A headline sample size should therefore be interpreted with attention to how the usable data were actually obtained.
A representation decision table
| Question | Useful representation |
|---|---|
| Compare separate categories | Bar chart |
| Show change over ordered time | Line graph |
| Show proportions of one whole | Pie chart |
| Show small numerical data in detail | Dot plot or stem-and-leaf |
| Show grouped continuous distribution | Histogram |
| Estimate median/quartiles from grouped totals | Cumulative-frequency graph |
| Compare median and spread compactly | Box plot |
| Retrieve exact values quickly | Table |
A data-quality workflow
- State the population or target group.
- Identify what variable is being measured.
- Check whether the sample is plausibly representative.
- Classify the variable as categorical, discrete or continuous.
- Tabulate frequencies accurately.
- Choose a representation that matches both the data and the question.
- Inspect scales, labels and grouping choices.
- Limit conclusions to what the collected evidence supports.
Common failure modes
| Error | Cause | Repair |
|---|---|---|
| Uses histogram for categories | Variable type ignored | Use bar chart for separate categories |
| Uses pie chart for unrelated totals | Whole-part condition missing | Confirm categories partition one whole |
| Creates overlapping class intervals | Boundary convention unclear | Use explicit inequalities |
| Generalises from convenience sample | Sampling bias ignored | Question representativeness |
| Trusts dramatic visual difference | Axis scale not inspected | Read numerical scale first |
| Chooses graph before identifying purpose | Representation treated as decoration | Ask what relationship must be visible |
Independent practice
- Classify “number of books read last month”.
- Classify “time taken to run 100 m”.
- Choose a display for monthly rainfall across one year.
- Choose a display for favourite sport categories.
- In a survey of 200 students, 50 choose option A. Find the pie-chart sector angle.
- Explain one reason why surveying only members of a chess club cannot represent all students’ interest in chess.
Explained answers
1. Discrete numerical.
2. Continuous numerical.
3. A line graph is suitable for ordered monthly change.
4. A bar chart is suitable for separate categories.
5. 50/200×360°=90°.
6. Chess-club members are already more likely than the general student population to be interested in chess, so the sample is biased.
Final thought
Statistical judgement begins before calculation. Good data work asks whether the right people were measured, whether the variable was classified accurately, whether the table preserves the observations and whether the display makes the intended relationship visible without distortion.
Collect carefully, classify clearly, tabulate faithfully and choose the representation that serves the question rather than decorating the answer.
Return to the Secondary Mathematics Hub.