Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

Primary 4 Science Learning Guide | Sampling, Representative Cases and Avoiding Cherry-Picking

Six plants sit on a bench.

Two are tall, three are medium-sized and one is unusually small.

A pupil wants to compare leaf size between two conditions. She chooses the largest plant from Condition A and the smallest plant from Condition B because the difference will be “easier to see”.

The measurement may be perfectly careful. The ruler may be correct. The table may be neat.

The evidence is still weak, because the selection was designed to favour a dramatic difference.

Science can be distorted before the first measurement is taken—simply by choosing which cases are allowed to become evidence.

This guide belongs to the Primary 4 Science Learning Hub. It develops a missing evidence-quality job: choosing plants, objects, locations, trials or observations in a way that represents the scientific question rather than the learner’s preferred result.

It does not teach formal statistical sampling, probability distributions or advanced randomisation. The P4 goal is simpler and very powerful: do not choose only the cases that make your story look good.

Quick Answer: The Fair-Sampling Loop

QUESTION → DEFINE WHAT COUNTS → IDENTIFY THE AVAILABLE SET → CHOOSE WITHOUT FAVOURING THE EXPECTED RESULT → RECORD NATURAL VARIATION → COMPARE → STATE THE LIMIT → TEST ANOTHER CASE

This is an eduKate teaching routine, not an official MOE examination formula.

Wait, What? A Perfect Measurement of a Bad Sample Is Still Weak Evidence

Suppose a child measures one leaf from a plant with extraordinary care.

The measurement may be accurate for that leaf.

But if the conclusion says:

“All leaves on this plant are 12 cm long,”

the evidence is too narrow.

The problem is not the ruler.

The problem is that one case has been asked to stand for a larger set without justification.

1. What is a sample?

For this P4 guide, a sample is the part of a larger set that the learner actually observes or measures.

Examples:

  • three leaves measured from a plant with many leaves;
  • four plants selected from ten available plants;
  • five locations observed along a garden trail;
  • three repeated readings from a longer data record;
  • two cups selected from a shelf of many cup designs;
  • one photograph chosen from a longer time sequence.

The sample should match the scientific question.

2. Whole-system question or sample question?

Question:

“What is the mass of this one object?”

Measure the object itself.

Question:

“What are leaf lengths like on this plant?”

One leaf may not be enough.

Question:

“Did our water cool over 15 minutes?”

You need the relevant temperature record, not a selection of only the readings that show the largest decrease.

Before sampling, decide whether the question concerns:

  • one specific item;
  • a repeated set of items;
  • a whole group;
  • a change through time;
  • a comparison between groups.

3. Representative does not mean average-looking by guess

Children sometimes hear “representative” and choose the object that looks most ordinary.

That may be reasonable in a narrow classroom demonstration, but it can also hide variation.

A representative sample is better understood as:

a set of cases chosen in a way that gives the important natural variation a fair chance to appear.

For Primary 4, this might mean:

  • do not choose only the tallest plants;
  • do not choose only the cleanest-looking leaves;
  • do not choose only readings that match the prediction;
  • do not choose only one sunny garden station when the question concerns the whole trail;
  • do not choose only the most dramatic photograph from a sequence.

4. Cherry-picking: choosing evidence after knowing what you want

Cherry-picking means selecting only the observations that support a preferred conclusion while ignoring relevant observations that do not.

Example:

A cooling experiment is repeated five times.

  • Trial 1: foam smaller decrease.
  • Trial 2: foam smaller decrease.
  • Trial 3: similar decreases.
  • Trial 4: cloth smaller decrease.
  • Trial 5: foam smaller decrease.

Weak response:

“Use Trials 1, 2 and 5 because they show the correct pattern.”

Stronger response:

keep all relevant trials, investigate Trial 4 and decide whether any exclusion has a documented method reason rather than a prediction reason.

5. Excluding a bad trial is not automatically cherry-picking

Suppose Trial 4 spilled half the water before the final temperature was measured.

That is a real method problem.

The group can:

  • keep the original record;
  • label the spill;
  • explain why the trial is not comparable;
  • repeat the trial properly;
  • avoid silently deleting it.

Scientific exclusion needs a reason connected to the method.

“It disagreed with our prediction” is not enough.

6. Natural variation is not a mistake

Living things vary.

Leaves on one plant can differ in:

  • length;
  • width;
  • position;
  • age;
  • orientation.

Different plants can vary even if they are the same type.

This variation should not be erased to make the table look neat.

If the question concerns a general pattern, the sample should acknowledge that living systems are not identical copies.

7. Sampling plants fairly

Suppose the class has eight comparable seedlings under one condition.

Weak selection:

choose only the healthiest three because they are easiest to measure.

This may exaggerate growth or hide variation.

Better:

agree on a selection method before inspecting which specimens give the desired answer.

For a simple classroom task, the teacher may pre-label plants and assign which ones are measured. The child then measures the assigned cases rather than choosing after seeing them.

Comparable does not mean identical

Plants should be similar enough that the intended comparison is meaningful.

If one plant begins twice as tall as another, comparing final heights alone may be misleading.

Record starting condition and decide what property should be compared:

  • final height;
  • change in height;
  • leaf number;
  • wilting score.

The sample and the property must work together.

8. Sampling leaves fairly

Question:

“How long are leaves on this plant?”

If the learner measures only the longest leaf, the answer describes the longest leaf.

If the learner measures only new tiny leaves, the answer describes young small leaves.

For a broader description, choose leaves using a rule decided before measuring.

Example teaching rule:

measure the first fully visible leaf at three pre-labelled positions.

The exact rule depends on the question. The important point is that it is set before the results are known.

9. Sampling objects fairly

Question:

“Do flexible solids keep a fixed volume?”

Choosing only one unusual object may create confusion about the broader property.

Use several suitable examples such as:

  • rubber eraser;
  • sponge;
  • soft clay piece if the lesson is designed carefully;
  • flexible plastic object.

Then ask which properties remain consistent and which appearances vary.

Do not choose only examples that behave dramatically.

10. Sampling locations fairly

A garden trail has:

  • two sunny stations;
  • three shaded stations;
  • one sheltered station.

A learner wants to describe “the garden temperature”.

One measurement in the hottest sunny location cannot automatically stand for the whole garden.

Possible better approaches:

  • narrow the question to that one location;
  • measure several pre-defined locations;
  • report each location separately;
  • avoid inventing one “garden temperature” if the site is variable.

11. Sampling time fairly

A time-lapse contains twelve images.

The learner selects the first and the most dramatic middle image, ignoring the later recovery.

That can distort the story.

When the question concerns change over time, the sample should preserve the relevant sequence or clearly state which times are being compared.

Batch 20’s Photographs, Video and Time-Lapse Observation guide develops the visual-record side of this problem.

12. Sampling repeated measurements fairly

A learner measures the same shadow five times:

  • 14 cm;
  • 15 cm;
  • 14 cm;
  • 14 cm;
  • 16 cm.

It is wrong to report only 14 cm because that value appears most convenient.

The whole relevant set should be considered.

At P4, the learner may describe:

“The repeated measurements ranged from 14 cm to 16 cm, with most values close to 14–15 cm.”

No advanced statistics are necessary.

13. Sampling observations before knowing the result

One strong way to reduce cherry-picking is to decide the selection rule before data collection.

Examples:

  • measure Plants A, C and E because the teacher assigned them before growth was seen;
  • record every tenth second rather than choosing interesting moments afterward;
  • photograph at 9 a.m. daily rather than only when the plant looks unusual;
  • measure the pre-labelled central leaf rather than the largest leaf.

This does not make the sample automatically perfect.

It reduces one source of bias: choosing after the preferred answer is visible.

14. Convenient samples

Convenience can be practical.

A child may choose the plant nearest the aisle because it is easy to reach safely.

That is not automatically wrong.

But ask:

“Could this convenient case be systematically different from the rest?”

If all aisle plants receive more light or are handled more often, convenience may affect representativeness.

The solution may be to narrow the conclusion to the sampled plants rather than claiming the whole group.

15. Dramatic examples are good for teaching, weak for prevalence claims

A very wilted plant can make the effect of root damage easy to discuss.

That is useful for teaching a mechanism.

But it does not show how common severe wilting is among all damaged plants.

One example can illustrate a possibility.

A sample is needed for a group-level description.

16. The “best-looking data” trap

A pupil repeats a shadow experiment and obtains one beautifully smooth set of values and one messier set.

They choose the smooth one for the report.

Ask:

Was the messy set invalid because of a documented method problem, or merely inconvenient?

If no method problem exists, both sets are relevant evidence.

Scientific reports should not become art competitions for the neatest graph.

17. Representative sampling and comparison groups

Suppose Condition A uses three medium-sized plants.

Condition B uses the three smallest plants.

Even if each condition is measured carefully, starting specimen differences can explain later differences.

Better design:

  • choose reasonably comparable starting specimens;
  • record the starting property;
  • use the same selection rule for both groups;
  • avoid choosing after knowing the final result.

18. Representative sampling and materials

Question:

“Are metals generally good conductors of heat?”

Testing one metal object can provide one example, not a universal survey of all metals.

At Primary 4, the syllabus can teach the general class property through appropriate examples and established knowledge.

A child’s small classroom sample should not be described as discovering the property of every possible metal.

This is an evidence-boundary lesson.

19. Sample size: more can help, but more is not magic

Using several cases can reduce dependence on one unusual individual.

Science Buddies’ procedure guidance, for example, recommends repeating experiments and gives the practical example of using several plants rather than relying on one plant. See Science Buddies: Experimental Procedure.

But more cases do not fix a biased selection rule.

Ten unusually large plants deliberately chosen because they support the prediction are still a biased sample.

Quantity and fairness both matter.

20. Why this guide does not prescribe a magic sample number

Different questions need different designs.

Three repeated readings can be useful for checking consistency.

Three plants may still be too few for a broad biological claim.

A single object is enough if the question concerns that specific object.

Therefore no universal P4 rule such as “always use five” is taught here.

The better questions are:

  • What population or set does the conclusion refer to?
  • Does the chosen sample give a fair view of the important variation?
  • Is the conclusion narrower than or equal to what was sampled?

21. Sampling and independent replication

Two groups can use the same method but different samples.

If their results differ, ask whether:

  • the method differed;
  • the samples differed naturally;
  • one group selected extreme cases;
  • the sample was too narrow;
  • the measured property was defined differently.

This connects directly with Replication, Reproducibility and Independent Group Checks.

22. Sampling and observer expectations

A learner expects damaged plants to wilt more.

They walk along the bench and choose the three most wilted damaged plants for photographs.

Now the selection reflects the expectation.

One repair is to pre-label the plants to be observed before the learner examines which ones look most dramatic.

Another is to include all relevant assigned specimens when the number is manageable.

This connects with the Batch 21 guide on Observer Expectations, Confirmation Bias and Independent Checks.

23. Research evidence and selective examples

Batch 20’s Researching with Books, Websites and Secondary Sources guide teaches source selection.

Sampling logic applies there too.

A pupil should not search until they find three webpages supporting the preferred answer and ignore suitable sources that challenge it.

Source selection can also be cherry-picked.

The learner should use a reasonable source rule and preserve relevant disagreement.

24. Original Sampling Casebook

Case 1 | Tallest vs shortest

Condition A gets tallest plant, Condition B gets shortest plant.

Weakness: starting specimen difference.

Repair: comparable selection rule and starting measurements.

Case 2 | Only the successful trials

Keep only repeats matching prediction.

Weakness: cherry-picking results.

Repair: keep all relevant trials; justify exclusions by method evidence.

Case 3 | One leaf represents all

One leaf measured carefully.

Weakness: sample too narrow for a whole-plant claim.

Repair: narrow conclusion or use a broader pre-defined sample.

Case 4 | Convenient sunny station

Only the hottest open station measured.

Weakness: not representative of mixed garden locations.

Repair: specify “sunny station temperature” or sample several defined station types.

Case 5 | Missing noon photograph

Only morning and afternoon frames exist.

Weakness: no noon evidence.

Repair: state the gap; do not invent a frame.

Case 6 | Best graph selected

One smooth trial shown, messy trial hidden.

Weakness: selective reporting unless messy trial had documented method failure.

Case 7 | Pre-labelled specimens

Teacher assigns Plants B, D and F before pupils inspect growth.

Strength: reduces result-driven selection.

Case 8 | Different starting sizes

Two groups compare final height only.

Weakness: starting conditions may explain final values.

Repair: compare starting and final values or change.

Case 9 | One unusual object

A very strange material example used for a broad class claim.

Weakness: exceptional case may not represent the intended group.

Case 10 | All assigned cases included

The class measures every plant in a small assigned set.

Strength: no selection among those cases after results are seen.

Case 11 | Extreme photograph

A dramatic frame chosen to represent an entire week.

Weakness: time sequence compressed selectively.

Repair: preserve regular captures or state why that moment was selected.

Case 12 | Product review search

Learner uses only pages praising an insulating cup.

Weakness: source sampling favours one outcome.

Repair: seek independent evidence and compare conditions.

25. The Sample-Selection Card

QuestionCheck
What larger set does my conclusion refer to?
What cases are actually available?
Did I choose the selection rule before seeing which result I preferred?
Are important starting differences recorded?
Am I keeping inconvenient but valid observations?
If I exclude something, is there a method reason?
Does natural variation need several cases?
Is my conclusion no broader than my sample?

26. A simple “choose before you know” strategy

Where suitable, decide the selection rule before inspecting the outcome.

Examples:

  • use pre-labelled plant numbers;
  • measure every second assigned object;
  • record at pre-set times;
  • use all valid trials within the planned set;
  • compare all sources meeting the stated research criteria.

This is not formal random sampling.

It is a child-friendly protection against choosing evidence after seeing which cases support the desired story.

27. When to use all cases instead of a sample

If the set is small and manageable, sometimes the simplest solution is to include all relevant cases.

Example:

There are four assigned cups.

Instead of choosing the two most different, measure all four if the question benefits and the procedure remains safe and manageable.

This removes the sample-selection problem within that small set.

It does not automatically make the conclusion universal beyond those four cups.

28. When sampling is unavoidable

A plant has forty leaves. Measuring every leaf may be unnecessary for a simple classroom question.

A trail has twenty observation points. Time may allow only five.

A data logger records one thousand values. A learner may need a summary rather than reciting all rows.

Then the sampling rule matters.

The learner should be able to explain why the selected cases are useful and what might be missed.

29. Sampling and ethical care

Do not damage many plants or living things merely to increase sample size.

If a question can be answered through non-destructive observation, teacher-provided records or a smaller ethically appropriate sample, choose the responsible route.

More data do not justify unnecessary harm.

30. Sampling and safety

Do not sample risky locations merely for representativeness.

A complete “garden sample” does not require stepping into traffic, deep water, unstable slopes or unsafe heat.

Safety is a legitimate constraint.

The correct response is to state the sampling boundary.

Example:

“Our observations cover the five safely accessible stations.”

31. Sampling and claim wording

Compare:

Too broad: “All plants grow 3 cm per week.”

Bounded: “The three assigned plants increased in measured height over this week, with changes between 2 cm and 4 cm.”

The second is less dramatic and more scientific.

A bounded claim makes the sample visible.

32. Original Practice Set

  1. What is a sample in this guide?
  2. Why can a perfectly measured leaf still be weak evidence for all leaves?
  3. What is cherry-picking?
  4. Is excluding a spilled trial always cherry-picking?
  5. Why should specimen selection sometimes happen before outcomes are known?
  6. What is natural variation?
  7. Why is choosing only the tallest plants risky?
  8. What should be recorded when groups start with different specimen sizes?
  9. Why is one sunny station not automatically “the garden temperature”?
  10. Why should a missing photograph remain missing?
  11. Do more cases always fix biased sampling?
  12. Why is there no magic P4 sample number?
  13. What is the difference between a dramatic example and a representative sample?
  14. How can source research also be cherry-picked?
  15. What is one child-friendly way to reduce selection bias?
  16. When might measuring all cases be easier than sampling?
  17. Why must exclusions be documented?
  18. How does sampling connect to independent replication?
  19. Why should safety limit sampling?
  20. Write one bounded conclusion from three plant measurements.

33. Practice Answers

1. The part of a larger set that is actually observed or measured.

2. It describes one leaf accurately but may not represent the variation among the others.

3. Selecting only evidence that supports a preferred result while ignoring relevant contrary evidence.

4. No. If the spill makes the trial scientifically incomparable, it can be excluded from a particular comparison while preserved and documented.

5. So the learner cannot choose only cases that already match the expected result.

6. Real individuals or repeated observations can differ even under similar conditions.

7. It may exaggerate the group’s typical height and bias comparisons.

8. Starting height or the relevant baseline, so final differences are interpreted correctly.

9. The garden contains different locations and conditions; one extreme location may not represent the whole.

10. Duplicating or inventing it would create evidence that was never captured.

11. No. Many deliberately selected extreme cases remain biased.

12. The appropriate number depends on the question, variation, available cases and conclusion.

13. A dramatic example shows a possibility clearly; a representative sample aims to reflect the relevant variation in the set.

14. A learner can select only websites supporting a preferred claim and ignore suitable contrary sources.

15. Pre-label or pre-assign cases before observing which ones support the prediction.

16. When the relevant set is small, safe and manageable.

17. So readers can distinguish a method-based exclusion from result-based cherry-picking.

18. Different groups may obtain different results because their samples differ naturally or were selected differently.

19. Representativeness never justifies entering hazardous locations or harming living things.

20. Example: “The three assigned plants increased by 2 cm, 3 cm and 4 cm over the observed week; these results describe the sampled plants and do not establish the growth of every plant.”

34. The Sampling Diagnostic

If the learner…Likely weak linkRepair
chooses extreme casesselection biaspredefine selection rule
hides disagreeing trialscherry-pickingkeep all valid evidence
generalises from one casesample boundarynarrow claim or broaden sample
treats variation as errornatural variationrecord several comparable cases
demands unsafe completenesssampling boundarystate safe accessible scope

35. A 40-Minute Sampling Lesson

Minutes 1–5: sort one-item vs group-level questions.

Minutes 6–10: identify biased plant selections.

Minutes 11–15: create a pre-selection rule.

Minutes 16–20: inspect natural variation in a safe prepared dataset.

Minutes 21–25: diagnose cherry-picked repeated trials.

Minutes 26–30: decide whether one exclusion is justified.

Minutes 31–35: write a bounded conclusion.

Minutes 36–40: transfer to sources, time-lapse or field locations.

36. What Parents and Tutors Can Ask

  • “What larger group are you trying to talk about?”
  • “Which cases did you actually measure?”
  • “How did you choose them?”
  • “Did you choose before or after seeing the result?”
  • “What variation did you keep?”
  • “Why was this case excluded?”
  • “Could the convenient examples be unusual?”
  • “Is your conclusion wider than your sample?”

37. Continue Batch 21

The Quiet Return

The ruler measures what is placed in front of it.

The harder question came earlier:

Why did this case get chosen?

When the learner can answer that question honestly—before seeing which answer would be most convenient—the evidence becomes more representative, the conclusion becomes more modest and the Science becomes stronger.