Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

Secondary 3 Mathematics Learning Guide | Probability and Statistical Reasoning

Probability asks what may happen. Statistics asks what the data already tells us. Strong mathematical judgement needs both. A probability can describe uncertainty without guaranteeing one outcome. An average can summarise a data set without describing its spread. A graph can reveal a pattern while its scale can also mislead. The mathematics becomes reliable only when the calculation and the interpretation agree.

This Secondary 3 Mathematics Learning Guide develops the Statistics and Probability strand through connected reasoning: collect and classify data, choose and interpret representations, compare centres and spreads, list outcomes, calculate probabilities, use tree or possibility diagrams, and decide when probabilities should be added or multiplied. It belongs to the Secondary Mathematics Hub.

The official 2027 SEC G3 Mathematics syllabus K310 includes data handling and analysis—tables, graphs, histograms with equal class intervals, stem-and-leaf diagrams, cumulative frequency, box-and-whisker plots, measures of centre and spread—and probability of single and simple combined events using possibility or tree diagrams, including mutually exclusive and independent events. Schools may sequence these ideas differently.

Use this route: diagnostic · representations · centre and spread · probability rules · tree diagrams · practice · answers. All examples are original teaching examples.

Statistics Starts With a Question, Not a Formula

A statistical calculation is meaningful only after the data and purpose are clear. Are we trying to describe a typical result? Compare two groups? Measure variability? Estimate a percentile? Detect a misleading graph? Different questions require different summaries.

The same data can support several valid representations. A stem-and-leaf diagram preserves individual values. A box plot compresses a data set into positional summaries. A cumulative frequency curve supports percentile estimates. A histogram emphasises frequency across intervals. The “best” representation depends on what needs to be seen.

A Six-Question Diagnostic

For the data 4, 6, 6, 8, 11, find the mean, median, mode and range. Then answer two probability questions: a fair six-sided die is rolled once—what is the probability of an even number? Two fair coins are tossed—what is the probability of two heads?

The answers are mean 7; median 6; mode 6; range 7; probability of an even die result 3/6 = 1/2; probability of two heads 1/4. These answers test different ideas. A student may calculate the mean correctly while confusing median position, or handle single-event probability while multiplying incorrectly in a combined event.

Tables and Graphs Are Representations of the Same Evidence

A table is often the most precise starting representation because categories and frequencies remain explicit. Graphs can make patterns easier to see, but every graph adds design choices: axis scale, interval width, category order and visual area.

A bar graph normally compares separate categories. A line graph often represents change across an ordered variable such as time. A pie chart represents parts of a whole. A histogram represents numerical intervals, so bars touch because the intervals form a continuous sequence.

The question “Which graph should I use?” should therefore be answered by identifying the variable type and the intended comparison before drawing anything.

Worked Example 1: A Misleading Vertical Scale

Two values are 82 and 86. A bar graph begins its vertical axis at 80 instead of 0. What effect can this have? The numerical difference is only 4 units, about 4.9% of 82. But if the displayed bars extend from 80, their visible heights are 2 and 6 units, so one appears three times as tall as the other.

The graph is not automatically “false” because the axis is truncated; some graphs deliberately zoom in to show small changes. The problem arises when the visual impression is interpreted without noticing the scale. Always read the axis labels before comparing apparent heights.

Stem-and-Leaf Diagrams Preserve Individual Values

For the data 12, 14, 17, 21, 21, 26, 30, a simple stem-and-leaf diagram can use tens digits as stems and units digits as leaves. The ordered display shows distribution shape while keeping every original observation recoverable.

A key is essential. “2 | 6 means 26” tells the reader what each digit position represents. Without a key, a leaf diagram using decimal data could be ambiguous.

Cumulative Frequency Is a Running Total

If frequencies across increasing score intervals are 4, 7, 10 and 5, the cumulative frequencies are 4, 11, 21 and 26. Each cumulative value includes all observations up to that boundary.

A cumulative frequency curve can be used to estimate the median, quartiles and percentiles. With 26 observations, the median lies around the 13th observation, while lower and upper quartiles correspond approximately to one quarter and three quarters of the ordered data. Exact school conventions for reading positions should follow the stated method and diagram.

Box Plots Compress Position, Not Every Detail

A box-and-whisker plot displays a five-number summary: minimum, lower quartile, median, upper quartile and maximum. It makes centre and spread easy to compare across groups.

However, a box plot does not show every individual observation. Two data sets can share the same five-number summary while differing internally. A compressed representation is useful because it hides detail—but that same compression limits what can be concluded.

Mean, Median and Mode Describe Different Ideas of Centre

The mean uses every value. The median identifies the middle position after ordering. The mode identifies the most frequent value or category. None is universally “the best average”.

For 4, 5, 5, 6, 30, the mean is 10 while the median and mode are both 5. The large value 30 pulls the mean upward. If the question asks for a typical central experience, the median may describe the group better than the mean. If every value contributes proportionally to a total, the mean may be the relevant calculation.

Worked Example 2: Compare Two Groups With the Same Mean

Group A scores are 48, 49, 50, 51, 52. Group B scores are 30, 40, 50, 60, 70. Both groups have mean 50.

Yet the ranges are very different: Group A has range 4, while Group B has range 40. The same centre can coexist with very different variability. A comparison based only on mean would therefore miss an important feature of the data.

Range and Interquartile Range Measure Spread Differently

Range is maximum minus minimum, so it depends entirely on two extreme values. Interquartile range, IQR = Q3 − Q1, describes the spread of the middle half of the ordered data.

Because the IQR ignores the outer quarters when measuring its width, it is less sensitive to extreme values than the range. This can make it useful when comparing distributions containing unusually high or low observations.

Standard Deviation Measures Typical Spread Around the Mean

Standard deviation uses all observations to quantify how dispersed values are around the mean. A smaller standard deviation indicates that values are more tightly clustered around the mean; a larger value indicates greater spread.

When comparing two data sets using mean and standard deviation, interpret both. A higher mean may indicate stronger average performance, while a lower standard deviation may indicate greater consistency. Neither automatically makes one group “better” without knowing the purpose of the comparison.

Worked Example 3: Grouped Mean

Suppose a grouped table has intervals 0–10, 10–20 and 20–30 with frequencies 3, 5 and 2. Using class midpoints 5, 15 and 25, the estimated mean is [3(5) + 5(15) + 2(25)] / 10 = 140/10 = 14.

This is an estimate because the exact observations inside each class are unknown. Treating every observation as if it equals its midpoint creates an approximation that becomes more reliable when intervals are sufficiently narrow and the within-class distribution is not extremely uneven.

Probability Is a Number Between 0 and 1

A probability of 0 represents an impossible event under the model. A probability of 1 represents a certain event. Values between them express different levels of chance.

For equally likely outcomes, probability = favourable outcomes / total outcomes. The phrase “equally likely” matters. Counting outcomes without checking equal likelihood can produce a wrong probability even when the fraction is simplified correctly.

Worked Example 4: One Die

A fair six-sided die is rolled once. Find the probability of obtaining a prime number. The possible outcomes are 1, 2, 3, 4, 5 and 6. The prime outcomes are 2, 3 and 5.

Therefore P(prime) = 3/6 = 1/2. The number 1 is not prime, so including it would change the event definition.

Mutually Exclusive Events Cannot Occur Together

On one roll of a die, “result is 2” and “result is 5” are mutually exclusive because both cannot happen in the same trial. For mutually exclusive events A and B, P(A or B) = P(A) + P(B).

But “result is even” and “result is greater than 3” are not mutually exclusive because 4 and 6 satisfy both. Adding the probabilities directly would count the overlap twice. In general, P(A ∪ B) = P(A) + P(B) − P(A ∩ B).

Worked Example 5: Addition With an Overlap

For a fair die, let A be “even” and B be “greater than 3”. Then A = {2, 4, 6}, B = {4, 5, 6}, and A ∩ B = {4, 6}.

P(A ∪ B) = 3/6 + 3/6 − 2/6 = 4/6 = 2/3. Listing the union {2, 4, 5, 6} gives the same result.

This is the probability version of the double-counting idea developed in Set Language, Venn Diagrams and Counting.

Independent Events Do Not Change Each Other’s Probabilities

Two fair coin tosses are independent under the standard model: the outcome of the first does not change the probability distribution of the second. For independent events A and B, P(A and B) = P(A)P(B).

Independence is different from mutual exclusivity. Independent events can occur together. Mutually exclusive events cannot. In fact, two mutually exclusive events with positive probability are not independent because occurrence of one makes the probability of the other zero.

Tree Diagrams Make Sequential Probability Visible

A tree diagram shows stages as branches. Probabilities along a complete path are multiplied because all events on that path must occur. Probabilities of separate favourable paths are added when any one of them gives the desired event.

Each set of branches leaving one node should total 1. This provides a simple check on the model before any path probabilities are calculated.

Worked Example 6: Two Independent Draws With Replacement

A bag contains 3 red and 2 blue counters. One counter is drawn, replaced, and a second counter is drawn. Find the probability of two red counters.

Because the first counter is replaced, each draw has P(red) = 3/5. Therefore P(RR) = (3/5)(3/5) = 9/25.

Replacement restores the original composition, so the second probability does not depend on the first colour. The independence here comes from the sampling rule, not merely from the fact that there are two draws.

Worked Example 7: Without Replacement

The same bag contains 3 red and 2 blue counters, but the first counter is not replaced. Find P(RR). The first red probability is 3/5. After a red counter is removed, 2 red counters remain among 4 total counters.

Therefore P(RR) = (3/5)(2/4) = 3/10. The second-stage probability changed, so the events are not independent.

Worked Example 8: Exactly One Red

Using the no-replacement model, exactly one red can occur by two paths: red then blue, or blue then red.

P(RB) = (3/5)(2/4) = 3/10. P(BR) = (2/5)(3/4) = 3/10. Add the separate favourable paths to get 3/5.

The phrase “exactly one” excludes RR and BB. Writing the event in words before calculating helps prevent including unwanted paths.

Possibility Diagrams Are Useful for Two Finite Outcomes

When two dice are rolled, a 6 × 6 possibility diagram displays 36 ordered pairs. This makes events such as “sum equals 7” easy to count: (1,6), (2,5), (3,4), (4,3), (5,2), (6,1), giving probability 6/36 = 1/6.

The diagram also exposes unequal numbers of ways to obtain different sums. A sum of 2 has one route, while a sum of 7 has six. Therefore the possible sums 2 through 12 are not equally likely.

Relative Frequency Connects Experiment to Probability

If an event occurs 47 times in 100 trials, its relative frequency is 0.47. This experimental proportion can be compared with a theoretical model, but it need not equal the theoretical probability exactly.

As the number of independent repeated trials grows under stable conditions, relative frequencies often settle nearer to the underlying probability. One small experiment does not “prove” the theoretical value wrong merely because its result differs.

Four Errors That Need Different Repairs

Mean used without spread: two groups are called equally consistent because their means match. Repair by comparing a spread measure as well.

Probability paths added when they should be multiplied: the student confuses “and” along one path with “or” across different paths. Repair by writing the event in words before the arithmetic.

Replacement ignored: second-stage probabilities remain unchanged even though an item was removed. Repair by recounting the sample space after the first outcome.

Graph interpreted by appearance alone: axis scale or interval choice is ignored. Repair by reading labels, numerical scale and units before describing the visual pattern.

Independent Practice

1. Find the mean, median, mode and range of 3, 5, 5, 7, 10.
2. Data set A is 9, 10, 10, 11, 10. Data set B is 2, 6, 10, 14, 18. Compare centre and spread.
3. A grouped table has intervals 0–20, 20–40, 40–60 with frequencies 2, 5, 3. Estimate the mean using midpoints.
4. A box plot has Q1 = 18 and Q3 = 31. Find the IQR.
5. Explain why two data sets with the same median need not have the same variability.

6. A fair die is rolled. Find P(number greater than 4).
7. A fair die is rolled. Find P(even or prime).
8. Two fair coins are tossed. Find P(exactly one head).
9. A bag contains 4 green and 1 yellow counter. Two draws are made with replacement. Find P(two yellow).
10. The same bag is used without replacement. Find P(green then yellow).

11. Two fair dice are rolled. Find P(sum = 9).
12. A tree branch has probabilities 0.7 and 0.3 from one node. Explain why this is a useful check.
13. Events A and B are mutually exclusive with P(A)=0.25 and P(B)=0.40. Find P(A ∪ B).
14. Independent events C and D have probabilities 0.6 and 0.5. Find P(C ∩ D).
15. In 200 trials an event occurs 74 times. Find the relative frequency and explain why it need not equal the theoretical probability exactly.

Explained Answers

1. Mean = 6; median = 5; mode = 5; range = 7.

2. Both have mean and median 10, but A has range 2 while B has range 16. B is much more spread out.

3. Midpoints are 10, 30 and 50. Estimated mean = [2(10)+5(30)+3(50)]/10 = 320/10 = 32.

4. IQR = 31 − 18 = 13.

5. Median records a central position only. Values on either side can be tightly clustered or widely dispersed while the middle value remains unchanged.

6. Outcomes 5 and 6 are favourable, so probability = 2/6 = 1/3.

7. Even = {2,4,6}; prime = {2,3,5}. Union = {2,3,4,5,6}, so probability = 5/6.

8. HT or TH gives probability 2/4 = 1/2.

9. P(YY) = (1/5)(1/5) = 1/25.

10. P(GY) = (4/5)(1/4) = 1/5.

11. Sum 9 occurs as (3,6), (4,5), (5,4), (6,3): 4/36 = 1/9.

12. Probabilities of all branches leaving one node should sum to 1; here 0.7 + 0.3 = 1.

13. Mutually exclusive events have no overlap, so 0.25 + 0.40 = 0.65.

14. Independence gives P(C ∩ D)=0.6×0.5=0.30.

15. Relative frequency = 74/200 = 0.37. Random variation in a finite experiment can make the observed proportion differ from a theoretical probability.

A Reliable Interpretation Routine

For statistics, identify the variable, representation, centre, spread and limitation of the conclusion. For probability, identify the sample space, event, dependence structure and whether the question uses “and”, “or”, “exactly” or “at least”.

Then check reasonableness. Probabilities must lie between 0 and 1. Frequencies cannot be negative. A subgroup cannot exceed the total group. A median must lie within the data range. These simple constraints catch many errors quickly.

Teacher and Parent Prompts

Ask “What does this number tell us, and what does it not tell us?” after a mean, probability or standard deviation is calculated. Ask the learner to describe a tree path in words before multiplying it. When a graph looks dramatic, ask for the actual numerical change and the axis scale.

For extension, give two data sets with the same mean but different spreads, or two events with the same individual probabilities but different dependence structures. The goal is to make interpretation decisions visible.

Continue the Secondary 3 Learning Route

Continue with Vectors and Geometric Relationships, Coordinate Geometry and Transformations, and Trigonometry, Bearings and Navigation.

Good statistical and probability reasoning does not stop at a correct calculation. It explains what the calculation means under the stated model and what conclusions the evidence actually supports. Return to the Secondary Mathematics Hub for the complete S1–S4 route.