Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

Secondary 1 Mathematics Learning Guide | Data, Averages, Statistical Representations and Probability

SECONDARY 1 MATHEMATICS LEARNING GUIDE · GUIDE 12

Data mathematics asks two questions at once: what do the numbers show, and what do they not show? A table can organise observations. A graph can make a pattern visible. An average can compress many values into one summary. A probability can describe how likely an outcome is under a model. Each representation is useful, but each also has boundaries.

A mean can be pulled by an extreme value. A bar chart can exaggerate a small difference if its vertical axis is truncated. A probability of 1/2 does not promise that exactly half of the next ten trials will produce the event. A pie chart cannot represent categories faithfully if the categories overlap.

This guide develops data handling, statistical representations, mean, median, mode, range and simple probability. It also trains the learner to inspect denominators, scales and assumptions before accepting a numerical conclusion.

Return to the Secondary Mathematics Hub. This guide connects to Percentages and Reverse Percentages and Coordinates, Linear Graphs and Relationships.

The current MOE Mathematics syllabuses include data handling and analysis, statistical representations and probability strands, with differences in exact depth and sequencing across subject levels. Use the MOE secondary syllabus directory as the official reference.

Navigate: the data question · tables and charts · mean, median and mode · spread and range · misleading graphs · probability · experimental probability · practice · answers.

1. Begin with what the data actually represent

Before calculating an average or drawing a chart, identify the observational unit and the variable. Are the data marks scored by students, travel times for journeys, temperatures on different days, or categories chosen in a survey?

The same numerical list can mean different things depending on the variable. A value of 12 could be twelve minutes, twelve people, twelve dollars or twelve points. The interpretation belongs to the data definition.

Population and sample

A population is the full group of interest. A sample is a subset observed in order to learn about that population. A sample can be useful, but its quality depends on how it is selected.

A survey of ten volunteers from one club cannot automatically represent every Secondary 1 student in Singapore. The arithmetic percentages may be correct for the sample while the broader conclusion is unsupported.

Ask what is missing

If a chart reports percentages but not the sample size, you may be unable to recover exact counts. If a mean is given without the distribution, you cannot reconstruct every original value.

2. Categorical and numerical data need different handling

Categorical data place observations into groups such as transport mode or favourite genre. Numerical data record quantities such as height, score or waiting time.

Some numerical-looking labels are actually categories. A bus number or postal code is not usually a quantity on which averaging makes sense.

Discrete and continuous numerical data

Discrete data take separate countable values, such as number of siblings. Continuous data can in principle take any value in an interval, such as mass or time, subject to measurement precision.

This distinction can influence how data are grouped and graphed later.

3. Frequency tables count how often values occur

A frequency table organises repeated observations. Suppose the values are 2, 3, 3, 4, 4, 4, 5.

ValueFrequency
21
32
43
51

The frequencies sum to 7, the number of observations.

Relative frequency

The relative frequency of value 4 is 3/7, approximately 42.9%. Relative frequency compares a count with the total number of observations.

The denominator is the whole data set represented by the table, unless the question defines another reference.

4. Bar charts compare categories using separated bars

A bar chart uses the height or length of each bar to represent a value or frequency for a category. The bars are separated because the categories are distinct.

The axis scale must be read before comparing bar heights. A difference of one centimetre on the page does not tell us the numerical difference without the scale.

Worked interpretation

Suppose a bar chart shows 18 students choosing option A, 24 choosing B and 12 choosing C. The total is 54. Option B represents 24/54 = 4/9 of the responses, approximately 44.4%.

The bar chart itself gives counts; the percentage is a derived comparison.

5. Line graphs emphasise change across an ordered variable

A line graph is useful when the horizontal variable has a meaningful order, often time. Connecting successive points can help show trend and change.

Do not assume every point between observations was actually measured. A connecting line may guide the eye rather than prove continuous behaviour.

Worked example

A temperature record is 25°C at 9:00, 27°C at 10:00 and 26°C at 11:00. The graph shows an increase of 2°C followed by a decrease of 1°C.

It does not tell us the exact temperature at 9:37 unless a model or additional measurement is supplied.

6. Pie charts represent parts of one whole

A full pie chart represents 360°, corresponding to 100% of the total. A category with fraction f of the whole receives angle 360f degrees.

Worked example

In an invented survey of 80 responses, 20 choose option A. The fraction is 20/80 = 1/4, so the sector angle is 360 × 1/4 = 90°.

Recover a count from a sector

If a sector angle is 72° in a pie chart representing 150 people, the fraction is 72/360 = 1/5. The corresponding count is 30 people.

Categories should form a coherent whole

If respondents can select multiple options, the category percentages may sum to more than 100%. A standard pie chart would then be inappropriate unless the categories are redefined as mutually exclusive parts of one total.

7. Pictograms require attention to the key

A pictogram uses symbols to represent counts. The key states the number represented by one full symbol.

Worked example

If one symbol represents 8 books, then 3.5 symbols represent 3.5 × 8 = 28 books, provided the pictogram convention permits a half-symbol to represent half the key value.

Do not count pictures without reading the key.

8. The mean balances total value across the number of observations

Mean = total of the values ÷ number of values.

Worked example

The values 6, 7, 9, 10 and 13 have total 45. The mean is 45 ÷ 5 = 9.

Recover a missing total

If 8 values have mean 12, their total is 8 × 12 = 96.

The relationship mean × number of observations = total is often more useful than treating the mean formula as one-way.

9. A new value changes the mean through the total

Suppose five values have mean 14. Their total is 70. A sixth value of 20 is added. The new total is 90, so the new mean is 90 ÷ 6 = 15.

Do not average the old mean 14 with the new value 20 to get 17. The old mean represents five observations, while the new value represents one.

Weighted thinking

The correct update respects how many observations each summary represents. This idea later connects to weighted averages and grouped data.

10. The median uses order, not total

The median is the middle value when data are arranged in order. With an odd number of observations, choose the central value. With an even number, take the mean of the two central values.

Worked example: odd count

For 3, 5, 8, 9, 12, the median is 8.

Worked example: even count

For 3, 5, 8, 9, 12, 20, the middle values are 8 and 9, so the median is 8.5.

The list must be ordered first. Choosing the middle entry from an unsorted list is not a median calculation.

11. The mode is the most frequent value or category

The mode is the value or category occurring most often. A data set can have one mode, more than one mode, or no mode if no value occurs more frequently than the others.

Worked example

For 2, 3, 3, 4, 4, 4, 6, the mode is 4.

Categorical mode

Unlike the mean and median, the mode can apply naturally to categorical data. If “bus” is the most common transport category, then bus is the mode.

12. Range is a simple measure of spread

Range = maximum − minimum. It describes the total span of the observed values but uses only the two endpoints.

Worked example

For 4, 7, 8, 9, 15, the range is 15 − 4 = 11.

Same mean, different spread

Data set A: 8, 9, 10, 11, 12 has mean 10 and range 4. Data set B: 0, 5, 10, 15, 20 also has mean 10 but range 20.

The same mean does not imply similar variability.

13. Mean, median and mode answer different questions

The mean uses every numerical value. The median identifies the middle ordered position. The mode identifies the most common value or category.

No single measure is universally “best”. The useful choice depends on the data and the question.

Extreme-value example

Consider 20, 21, 22, 23 and 100. The mean is 37.2, while the median is 22. The extreme value 100 pulls the mean upward.

If the purpose is to describe a typical central value for the four values clustered near the low twenties, the median may communicate that centre more clearly. But the mean remains a correct arithmetic summary.

14. Frequency tables can produce means efficiently

When values repeat, multiply each value by its frequency, add the products and divide by total frequency.

Value xFrequency ffx
122
236
414

Total frequency = 6 and Σfx = 12. Mean = 12/6 = 2.

Why this works

The frequency column tells us how many times each value appears. Multiplication reconstructs the contribution of those repeated values to the total.

15. Graph design can change perception without changing the data

Suppose two values are 98 and 102. A bar chart with vertical axis from 0 to 110 shows a modest difference. A chart starting at 97 can make the bars look dramatically different.

The numerical difference remains 4 in both charts. The truncated scale changes visual emphasis.

Does a truncated axis automatically make a graph dishonest?

No. A restricted scale can help reveal small changes, especially in scientific or financial contexts. The problem is when the scale is hidden, unclear or used to support an exaggerated interpretation.

Read labels and intervals

Check whether the axis begins at zero, whether intervals are equal, whether units change, and whether any break in the axis is marked.

16. Three-dimensional decoration can distort a chart

In a 3D-style bar or pie chart, perspective may make front regions look larger or back regions smaller. Decorative volume can distract from the one-dimensional bar height or sector angle that encodes the data.

For accurate reading, follow the axis or stated values rather than judging apparent area or volume.

Design principle

A statistical display should make the encoded quantity easy to compare. Decoration is secondary to legibility and truthful scaling.

17. Correlation is not automatically causation

If two measured quantities tend to rise together, that is an association. It does not by itself establish that one causes the other.

A third variable may influence both, the direction of influence may be different from assumed, or the pattern may be coincidental.

Mathematical reading

When the question asks only for a trend in the plotted data, describe the trend. Do not add a causal story unless the evidence or experimental design supports it.

18. Probability measures chance on a scale from 0 to 1

A probability of 0 represents an impossible event under the model. A probability of 1 represents a certain event. Values between them express intermediate likelihood.

Probabilities can also be written as fractions, decimals or percentages.

Equally likely outcomes

When all outcomes in a finite sample space are equally likely, P(event) = number of favourable outcomes ÷ total number of outcomes.

This formula should not be used when outcomes are not equally likely unless the probabilities are otherwise accounted for.

19. Sample spaces make possible outcomes visible

For one fair six-sided die, the sample space is {1,2,3,4,5,6}. The probability of rolling an even number is 3/6 = 1/2.

Complement

If event A has probability P(A), then the probability that A does not occur is 1 − P(A).

For the fair die, P(not even) = 1 − 1/2 = 1/2.

Impossible values

A probability cannot be −0.2 or 1.3 in ordinary probability. Such an answer reveals an error in the model or calculation.

20. Two-step experiments need all ordered outcomes

If a fair coin is tossed twice, the ordered outcomes are HH, HT, TH and TT. There are four equally likely outcomes under the ideal model.

The probability of exactly one head is 2/4 = 1/2.

HT and TH are different ordered outcomes

They both contain one head, but they occur in different orders. Collapsing them too early can undercount the sample space.

Tree diagrams and tables

A tree diagram or two-way table can make multi-stage outcomes visible. Use the representation that helps you count without omission or duplication.

21. “At least one” often becomes easier through the complement

When a fair coin is tossed three times, finding the probability of at least one head directly requires counting many outcomes. The complement is no heads, which means TTT.

P(no heads) = 1/8, so P(at least one head) = 1 − 1/8 = 7/8.

Complement reasoning is useful when the opposite event is simpler to count.

22. Independence is a condition, not a default assumption

Two events are independent when the occurrence of one does not change the probability of the other under the model.

Repeated fair coin tosses are commonly modelled as independent. Drawing objects without replacement from a bag is generally dependent because the contents change after a draw.

Course boundary

Formal multiplication rules for independent events may belong later in some programmes. The key Secondary 1 habit is to ask whether the first outcome changes the conditions for the second.

23. Experimental probability comes from observed frequency

If an event occurs 37 times in 100 trials, its experimental probability is 37/100 = 0.37.

This observed proportion may differ from a theoretical probability, especially in a small number of trials.

Long-run reasoning

Under a stable random model, relative frequency may tend to settle near the theoretical probability as the number of trials becomes large. This does not mean it must move closer on every additional trial.

Random variation remains present.

24. Expected frequency is a model-based prediction

If the probability of an event is 0.3 and an experiment is repeated 200 times under stable conditions, the expected frequency is 0.3 × 200 = 60.

This does not guarantee exactly 60 occurrences. It is the long-run average count predicted by the model.

Whole-count interpretation

An expected frequency can be non-integer, such as 12.5. That does not predict half an actual event in one experiment. It is an average expectation across repeated comparable experiments.

25. Averages and probabilities can both hide structure

A mean of 70 marks does not show whether every student scored near 70 or whether scores were widely spread. A probability of 0.6 does not show the sequence in which successes and failures will occur.

Summary numbers are useful because they compress information. Their limitation is the same compression.

Ask for the underlying distribution when it matters

When comparing two groups, inspect sample sizes, spread and data shape where available. Do not demand information the question does not provide, but do not pretend a summary contains it.

26. Common data and probability errors

ErrorLikely issueRepair prompt
Averages percentages from unequal groups directlyDifferent denominators ignoredHow many observations does each percentage represent?
Median chosen from unsorted dataOrder requirement missedWhere is the middle after sorting?
Mean updated by averaging old mean with new valueOld mean’s weight ignoredWhat total did the old mean represent?
Bar heights compared without scaleAxis ignoredWhat numerical interval does one grid step represent?
Pie chart used for overlapping categoriesWhole-part structure invalidCan one observation belong to several sectors?
Probability exceeds 1Sample space or denominator wrongCan the favourable outcomes exceed all possible outcomes?

27. Practice laboratory

  1. Find the mean of 6, 8, 9, 10 and 12.
  2. Find the median of 4, 11, 7, 5, 20.
  3. Find the median of 3, 4, 7, 9, 12, 18.
  4. Find the mode and range of 2, 5, 5, 6, 8, 8, 8, 12.
  5. Six values have mean 14. Find their total.
  6. Five values have mean 12. A sixth value of 18 is added. Find the new mean.
  7. A pie chart sector represents 30 out of 120 students. Find its angle.
  8. A 72° pie sector represents how many people out of 200?
  9. A fair six-sided die is rolled. Find P(number greater than 4).
  10. A fair coin is tossed twice. Find P(exactly two heads).
  11. A fair coin is tossed twice. Find P(at least one head).
  12. An event occurs 46 times in 80 trials. Find its experimental probability.
  13. A model gives an event probability 0.25. Find the expected frequency in 160 trials.
  14. Explain why a mean of 50 does not determine the range.
  15. Explain one way a truncated graph axis can exaggerate visual differences.
  16. A survey allows each participant to choose several hobbies. Explain why a standard pie chart of hobby percentages may be inappropriate.

28. Explained answers

1. Total = 45. Mean = 45/5 = 9.

2. Order: 4,5,7,11,20. Median = 7.

3. Middle values are 7 and 9. Median = 8.

4. Mode = 8. Range = 12 − 2 = 10.

5. Total = 6 × 14 = 84.

6. Old total = 5 × 12 = 60. New total = 78. New mean = 78/6 = 13.

7. 30/120 = 1/4, so angle = 90°.

8. 72/360 = 1/5. Count = 200/5 = 40.

9. Outcomes 5 and 6 are favourable. Probability = 2/6 = 1/3.

10. Only HH gives two heads, so probability = 1/4.

11. Complement is TT with probability 1/4, so answer = 3/4.

12. 46/80 = 0.575.

13. 0.25 × 160 = 40.

14. Many different data sets can have the same total and mean but different minimum and maximum values, so the range is not fixed.

15. Starting the axis close to the data values can make a small numerical difference occupy a large visual fraction of the chart height.

16. Pie sectors should partition one whole into non-overlapping categories. Multiple hobby selections can make category percentages sum beyond 100%.

29. Complete mixed problem: compare two groups carefully

Problem: Group A contains 10 values with mean 18. Group B contains 30 values with mean 22. Find the combined mean. Explain why averaging 18 and 22 directly is wrong.

Group A total = 10 × 18 = 180. Group B total = 30 × 22 = 660. Combined total = 840 across 40 values.

Combined mean = 840 ÷ 40 = 21.

The simple average of 18 and 22 is 20, but that gives the two groups equal weight. Group B contains three times as many observations, so it must contribute three times as much weight to the combined mean.

Now add spread

If Group A has range 4 and Group B range 20, the combined mean still does not tell us which group is more consistent. The spread information provides another dimension of comparison.

30. Teaching data mathematics as evidence reading

Begin with one small data set and ask several questions: mean, median, mode, range and an appropriate graph. Then change one value to an extreme number and ask which summaries move most.

Next show the same counts using two different axis scales. Ask what changed in the data and what changed only in the visual emphasis.

For probability, list the sample space before using a formula. Then alter the experiment from replacement to no replacement and ask whether the conditions remain the same.

A strong data habit

Every summary should be read with its denominator, scale, sample and model. This prevents correct arithmetic from becoming an unsupported claim.

31. Questions students often ask

Which average should I use?

Follow the question when it specifies one. When choosing a summary, consider whether extreme values, repeated categories or ordered position matter.

Can a data set have two modes?

Yes. If two values share the highest frequency, the data are bimodal. More generally, multiple modes are possible.

Does probability 0.7 mean seven successes in every ten trials?

No. It describes a long-run likelihood under the model, not a guaranteed short-run pattern.

Why can experimental probability differ from theoretical probability?

Random variation affects finite trials. With more trials under stable conditions, relative frequency may become more stable, but exact equality is not guaranteed.

What is the best data check?

Check the total frequency, denominator, axis scale, units, summary definition and whether the conclusion says more than the evidence supports.

32. Return path

Data and probability complete this Secondary 1 learning-guide batch by bringing together percentages, graphs and numerical judgement. Revisit Percentages and Reverse Percentages when relative frequency and part-to-whole comparisons are unstable. Revisit Coordinates, Linear Graphs and Relationships when the main difficulty is scale or graph interpretation.

Use Read the Question Before Choosing a Method whenever a statistics question contains several summaries and the learner must decide which one answers the actual claim.

Sources and learning boundaries

Official curriculum reference: MOE Secondary Syllabus Directory. The current G2/G3 syllabus includes data handling, statistical representations, measures of central tendency and probability strands; exact sequencing and depth vary.

Supplementary foundational reading: OpenStax, introductory statistics and sampling, plus its probability sections. All survey counts, group values and experiments above are original constructed teaching examples.

Editorial approach: Wintour House V1.0 · CivDJ · eduKate Publishing. Define the data, preserve the denominator and scale, choose the summary, test the interpretation and return the claim to the evidence.

Return to the Secondary Mathematics Hub →