A statistical diagram can only be as useful as the data structure underneath it. Before a bar is drawn or a mean is calculated, someone has decided what to measure, how to classify observations, which categories or intervals to use and which representation will make the important pattern visible.
This Secondary 3 Mathematics Learning Guide develops data collection, classification, tabulation and representation choice. It complements the existing guides on statistical summaries by focusing on what happens before a histogram, box plot or standard deviation is interpreted.
Official scope: the 2027 SEC G3 Mathematics syllabus K310 explicitly includes simple concepts in collecting, classifying and tabulating data, analysis and interpretation of several statistical representations, their purposes and uses, advantages and disadvantages, simple inference and explanations of misleading diagrams. This guide is G3-oriented; schools may introduce individual representation types earlier.
Data route: define the question → identify the variable → decide how observations will be obtained → classify consistently → tabulate accurately → choose a representation that matches the variable and purpose → interpret only what the data supports.
For distribution summaries, continue to Histograms, Cumulative Frequency, Box Plots and Standard Deviation.
Begin With a Question That Can Be Answered by Data
“What do students think?” is too vague for a mathematical investigation. “How many minutes do students in this class report spending on the journey to school?” identifies a measurable variable and a group to be observed.
A good data question specifies enough structure that two people could collect comparable observations. If the definition changes from person to person, the final table may contain numbers that do not mean the same thing.
Categorical and Numerical Variables Need Different Representations
A categorical variable records labels such as transport mode: walk, bus, train or car. A numerical variable records numbers such as journey time, number of siblings or temperature.
Bar graphs and pie charts are often suited to categorical frequencies. Histograms and stem-and-leaf diagrams are designed for numerical distributions. Choosing the wrong representation can conceal the structure or imply relationships that are not present.
Discrete and Continuous Numerical Data
Discrete data arise from countable values such as number of books borrowed: 0,1,2,3,… . Continuous data arise from measurements that can, in principle, take any value within an interval, such as height, mass or time.
The distinction helps with grouping. A class such as 150≤h<155 cm describes a measurement interval. A frequency table of exact sibling counts uses separate integer values rather than pretending that 1.5 siblings is an ordinary observation.
Worked Example 1: Classify the Variables
Classify: transport mode to school, number of pets, journey time, favourite school subject.
Transport mode and favourite subject are categorical. Number of pets is numerical discrete. Journey time is numerical continuous when measured rather than rounded into fixed categories.
The purpose of classification is practical: it narrows the sensible table and graph choices before any plotting begins.
Collection Method Must Match the Variable
Some data are obtained by measurement, some by observation and some by asking people to report information. Each method has limitations.
Measuring journey time with a clock may be more precise than asking someone to remember yesterday’s journey. Observing transport arrivals can record actual events. A survey is appropriate for preferences that cannot be measured physically.
For extension, also ask whether the observed group fairly represents the population about which a conclusion is made. K310 names the simpler collection/classification task; representative sampling is a useful statistical habit because conclusions should not silently reach beyond the evidence collected.
Worked Example 2: Improve a Survey Question
A survey asks: “Don’t you agree that our school canteen is excellent?”
The wording encourages agreement and mixes opinion with a strong positive description. A more neutral question could be: “How would you rate the school canteen overall?” followed by clearly defined response options.
Changing question wording can change responses. A numerical chart cannot remove bias introduced before the data were collected.
Classifications Must Be Mutually Clear
Categories should tell the recorder where each observation belongs. If age groups are written “10–12” and “12–14”, the value 12 appears to belong in both.
For continuous data, interval notation such as 10≤x<12 and 12≤x<14 removes the overlap. For categorical data, define ambiguous labels when necessary.
Worked Example 3: Repair Overlapping Classes
A table groups times as 0–10 min, 10–20 min and 20–30 min. If boundary values are possible, write the classes explicitly as 0≤t<10, 10≤t<20, 20≤t<30, with a separate rule for a possible endpoint of exactly 30 if the study includes it.
The repair is not cosmetic. It ensures each observation is counted exactly once.
A Frequency Table Is a Counting Machine
A raw list preserves individual observations but can be difficult to scan. A frequency table compresses repeated values or classes by recording how often each occurs.
Always check that the frequencies add to the number of observations. This simple total is one of the strongest ways to catch a missing or duplicated count.
Worked Example 4: Raw Data to Frequency Table
The reported numbers of books read in a month are: 2,1,3,2,0,4,2,1,3,2,4,1.
| Books | Frequency |
|---|---|
| 0 | 1 |
| 1 | 3 |
| 2 | 4 |
| 3 | 2 |
| 4 | 2 |
The frequencies total 1+3+4+2+2=12, matching the raw list. The modal value is two books because it occurs four times.
Grouped Tables Trade Detail for Compression
Grouping measurements into intervals makes a large distribution easier to inspect but removes the exact individual values. If seven observations lie in 20≤x<30, the table does not tell us whether those observations cluster near 20, near 30 or are spread evenly.
This information loss is why grouped means and standard deviations are estimates when class midpoints represent the hidden observations.
Worked Example 5: Group Continuous Data
Journey times in minutes are 8,12,17,19,21,24,25,28,31,33,37,39.
| Time interval | Frequency |
|---|---|
| 0≤t<10 | 1 |
| 10≤t<20 | 3 |
| 20≤t<30 | 4 |
| 30≤t<40 | 4 |
The total remains 12. The grouped table makes interval concentration visible but no longer shows, for example, that 21 and 28 are different observations inside the same class.
Choose a Bar Graph for Separate Categories
A bar graph is useful for comparing frequencies or values across separate categories. The spaces between bars reinforce that the categories are distinct rather than adjacent numerical intervals.
Category order can often be rearranged without changing meaning, although an ordered category scale may have a natural sequence.
Choose a Pie Chart for Part-to-Whole Composition
A pie chart emphasises how a total is divided among categories. Sector angle = category frequency / total × 360°.
It is less useful when many categories are close in size or when precise comparison matters. A bar graph may make small differences easier to see.
Worked Example 6: Convert Frequencies to Pie-Chart Angles
Transport modes among 40 students: walk 8, bus 12, train 16, car 4.
Walk angle=8/40×360=72°. Bus=108°. Train=144°. Car=36°.
The four angles total 360°, a direct construction check.
Choose a Line Graph When Ordered Position or Time Matters
A line graph can emphasise change across an ordered horizontal variable such as time. Joining points suggests continuity or progression between neighbouring x-values, so it should not be used casually for unrelated categories.
For monthly temperature readings, chronological order matters. For favourite colours, connecting red to blue with a line could imply a progression that has no defined meaning.
Histograms Are for Numerical Intervals
In K310, histograms are specified with equal class intervals. The bars touch because the horizontal axis represents adjacent numerical intervals rather than separate categories.
A bar graph of transport modes and a histogram of journey times may both contain rectangles, but their axes encode different structures. Representation choice begins with the variable, not the appearance of the finished graph.
Stem-and-Leaf Diagrams Preserve Individual Values
A stem-and-leaf diagram can display distribution shape while retaining individual numerical observations. This is useful for small-to-moderate data sets where both the overall pattern and exact values matter.
Always include a key, such as “2|7 means 27 minutes”, so the place value is unambiguous.
Dot Diagrams Make Repetition Visible
Dot diagrams place one mark for each observation above its value. They are useful for seeing clusters, gaps and repeated small data sets without compressing the observations into intervals.
As data sets become larger or measurement values become dense, another representation may communicate the distribution more efficiently.
Cumulative Frequency and Box Plots Answer Position Questions
Cumulative frequency is useful for estimating medians, quartiles and percentiles from grouped numerical data. A box-and-whisker plot compresses minimum, lower quartile, median, upper quartile and maximum into a compact comparison.
These representations deliberately discard or compress some detail. They are powerful when the question concerns centre and spread, but they are not the best choice when exact individual frequencies are required.
Worked Example 7: Match the Representation to the Question
Choose a useful representation for each task.
- Compare frequencies of four transport categories → bar graph.
- Show the percentage share of a total budget across five categories → pie chart, if the part-to-whole emphasis is the main purpose.
- Show how a quantity changes each month → line graph.
- Show a continuous numerical distribution across equal intervals → histogram.
- Compare median and IQR of two classes compactly → box plots.
Other representations may also be defensible if they answer the intended question clearly. Representation choice is a reasoned decision, not a one-to-one vocabulary test.
A Misleading Diagram Can Use Correct Numbers
A graph can contain genuine data and still exaggerate or hide differences through its scale, axis starting point, category width or selective time window.
For example, values 98,100 and 102 shown on a vertical axis from 97 to 103 can look dramatically different. The same values on an axis from 0 to 120 look much closer. Neither scale automatically changes the data, but each visual frame influences perception.
Worked Example 8: Truncated Axis
Two monthly values are 50 and 52. A chart begins its vertical axis at 49 rather than zero. The second bar may appear three times as tall above the displayed baseline even though the actual increase is only 2/50=4%.
The correct explanation is not simply “axes must always start at zero”. Some graphs legitimately use restricted ranges. The important task is to notice the scale and avoid interpreting visual height as though the baseline were zero.
Pictograms Depend on the Key
In a pictogram, one symbol may represent more than one observation. A half-symbol may therefore represent half of the key quantity, if that convention is stated.
Do not count pictures without reading the key. The pictorial appearance is a representation of frequency, not the frequency itself.
Worked Example 9: Pictogram Reading
If one symbol represents four students, 3½ symbols represent 14 students. If another category has two symbols, it represents eight students. The difference is six students, not 1½ students.
Inference Must Stay Within the Evidence
A class survey can describe that class. It does not automatically describe every student in Singapore. A line graph showing two quantities rising together does not by itself prove one causes the other.
K310 asks students to draw simple inference from statistical diagrams. A careful inference says what the displayed data supports and stops before adding conclusions not established by the representation.
Worked Example 10: Strong and Weak Claims
A school club records weekly attendance of 22,25,28,29 and 33 across five weeks.
A supported statement is: attendance increased overall across the five recorded weeks. A stronger statement such as “the club will definitely have more than 33 students next week” is not supported by these five observations alone.
Tables Are Often Better Than Graphs for Exact Lookup
A graph emphasises pattern; a table often preserves exact listed values. If a parent wants the exact fee for each category or a scientist needs exact recorded measurements, a table may serve the lookup purpose better.
There is no rule that every data set must become a graph. Representation should answer the task.
Worked Example 11: One Data Set, Two Useful Representations
A table lists daily temperatures 29,31,30,32,33,31,30°C for seven days. The table makes exact values easy to read. A line graph makes the day-to-day movement easier to see.
Neither is universally better. The table answers exact lookup; the line graph emphasises change over ordered time.
Worked Example 12: Reclassifying Can Change the Visual Story
Suppose numerical scores are grouped into 0–20,20–40,40–60,60–80,80–100. A second analyst uses 0–50 and 50–100. Both groupings contain the same raw data, but the coarser two-class table hides much more detail.
Grouping is therefore a modelling choice. Wider classes simplify the display while losing local structure. Narrower classes retain more shape but can make small data sets look noisy.
A Reliable Data-Handling Checklist
- What population or group does the question concern?
- What exactly is the variable?
- Is it categorical, discrete numerical or continuous numerical?
- How were observations obtained?
- Does every observation fit exactly one category or class?
- Do frequencies total the number of observations?
- Which representation best answers the intended question?
- Are axes, keys, units and scales clear?
- What information is lost by grouping or summarising?
- Does the final inference stay within the evidence?
Common Errors and Their First Repair
Choosing a graph by appearance: classify the variable and purpose first.
Overlapping intervals: write explicit lower and upper boundaries.
Frequencies do not match the data count: return to the raw list and tally systematically.
Treating a histogram like a categorical bar chart: remember that its horizontal axis contains adjacent numerical intervals.
Ignoring a pictogram key: convert symbols into frequencies before comparing.
Making a population-wide claim from a narrow group: state exactly which observations the evidence describes.
Independent Practice
1. Classify each variable: shoe size, favourite fruit, height, number of siblings.
2. Explain why “10–15” and “15–20” are ambiguous classes if 15 can occur, and rewrite them precisely.
3. Raw data are 1,2,2,3,4,2,1,0,3,2. Construct a frequency table for 0 to 4 and check the total.
4. Which is more suitable for comparing frequencies of five unrelated categories: a bar graph or line graph? Explain.
5. Thirty students choose clubs A,B,C,D with frequencies 6,9,12,3. Find the four pie-chart angles.
6. Explain one advantage of a stem-and-leaf diagram over a grouped histogram for a small numerical data set.
7. Explain one advantage of a box plot when comparing two large distributions.
8. A pictogram key says one symbol represents 6 students. How many students do 4½ symbols represent?
9. A vertical axis begins at 98 and two bars represent 100 and 102. Explain how the display could exaggerate the visual difference.
10. A survey asks only members of a school basketball team whether students enjoy basketball. Explain why a conclusion about the whole school should be cautious.
11. Choose a useful representation for monthly rainfall across one year and justify the ordered horizontal axis.
12. A grouped table contains frequencies 3,7,8,5,2. What simple check should be performed before any mean or graph is produced, and what total should result?
Explained Answers
1. Shoe size: numerical discrete under ordinary recorded sizes; favourite fruit: categorical; height: numerical continuous; number of siblings: numerical discrete.
2. Fifteen appears in both written ranges. One repair is 10≤x<15 and 15≤x<20.
3. Frequencies are 0→1, 1→2, 2→4, 3→2, 4→1. Total=10.
4. Bar graph; separate categories do not have a meaningful continuous order requiring connecting segments.
5. A:72°, B:108°, C:144°, D:36°.
6. It can preserve the individual values while still showing distribution shape.
7. It compactly compares median, quartiles and spread without plotting every observation.
8. 4.5×6=27 students.
9. The displayed baseline is only two units below 100, so the two-unit difference occupies a large proportion of the visible bar height even though 102 is only 2% above 100.
10. The selected group has a direct connection to basketball and may not represent preferences across the whole school.
11. A line graph can show change through the ordered months; a bar graph could also be defensible when precise monthly comparison is the main purpose.
12. Add the frequencies to verify the observation count. The total is 25.
How to Know the Data Skill Has Transferred
Do not always ask students to draw a graph after being told its name. Give a question and a raw data structure, then ask them to choose and justify the representation. Change from categorical to continuous data, or ask for exact lookup instead of pattern recognition. The student should alter the representation because the purpose changed.
A stronger task asks what the chosen representation hides. A box plot hides individual values. A grouped histogram hides exact positions inside each interval. A pie chart can make close categories hard to compare. Statistical maturity includes knowing both what a representation reveals and what it compresses.
Continue the Secondary 3 Learning Route
Continue with Fractional Equations, Restrictions and Equation Recovery, Financial Mathematics: Taxation, Instalments, Bills and Currency Exchange, and Map Scales, Floor Plans and Scale-Area Reasoning.
A statistical representation is secure when the collection method, classification, table, graph and conclusion all describe the same data without silently changing its meaning. Return to the Secondary Mathematics Hub.