Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

How to Decide Which PSLE Science Investigation Gives Stronger Evidence for a Claim

Wait, What? More Equipment Does Not Automatically Mean Better Evidence

Two students investigate the same scientific claim.

Student A uses six pieces of apparatus, records twenty numbers and repeats the procedure five times. Student B uses a simpler setup, measures one directly relevant outcome and keeps the comparison fair.

Which investigation gives stronger evidence?

You cannot decide by counting equipment, tables or repeated readings.

Evidence is strong when the investigation is well matched to the claim it is trying to test.

More data can strengthen a good design. More data cannot rescue a comparison that measures the wrong outcome, changes several conditions at once, uses the wrong specimens for the claim or collects measurements that do not answer the scientific question.

Quick Answer

When two or more PSLE Science investigations address the same claim, compare them in this order:

STATE THE CLAIM → IDENTIFY THE CHANGED CONDITION → IDENTIFY THE MEASURED OUTCOME → CHECK WHETHER THE COMPARISON IS FAIR → CHECK WHETHER THE MEASUREMENT DIRECTLY ANSWERS THE CLAIM → CHECK REPEATS AND VARIATION → CHECK SPECIMENS AND GENERALISATION → CHECK THE TESTED RANGE → CHECK CONTROL OR REFERENCE LOGIC → CHECK MEASUREMENT QUALITY → CHECK WHAT EACH CONCLUSION CAN ACTUALLY SUPPORT.

The stronger investigation is the one that gives the more defensible evidence for this claim. There is no single universal ranking such as “more repeats always wins” or “more apparatus means more scientific”.

The Exact PSLE Science Learning Job This Guide Owns

This guide owns one PSLE Science learner job: how a Primary 5 or Primary 6 learner compares two or more investigations that address the same scientific claim and decides which provides stronger evidence by matching design quality to the claim.

It does not replace the fair-test owner, the method-evaluation owner, the repeated-trials owner or the generalisation owner. Those pages teach individual evidence problems. This page owns the higher-level comparison:

Given two different investigations, which one is better evidence for the claim—and why?

The Current 2026 PSLE Science Frame

For examination from 2026, Standard PSLE Science assesses attainment in the 2023 Primary Science syllabus. The official assessment objectives include knowledge with understanding, application of scientific facts, concepts and principles, and scientific inquiry involving prediction or hypothesis, interpretation and analysis, evaluation of observations, information and methods, and communication of explanations and reasoning.

This means learners are expected not only to collect or read data, but also to evaluate whether the information and methods justify a scientific conclusion.

Evidence Strength Is Claim-Dependent

The same investigation can be strong evidence for one claim and weak evidence for another.

Suppose a student measures the final temperature of two cups after ten minutes.

  • For the claim “Cup A has a higher temperature after ten minutes,” the endpoint measurements may be excellent evidence.
  • For the claim “Cup A warmed faster during the first three minutes,” the same endpoint data are weak because they do not show what happened during the first three minutes.

Before ranking investigations, state the claim precisely.

The Eight Main Evidence-Strength Questions

QuestionWhy it matters
Does the method test the right claim?An investigation can be well executed but irrelevant to the question.
Is the comparison fair enough for the claim?Several changing conditions make causal attribution weaker.
Is the measured outcome relevant?A precise measurement of the wrong quantity is still the wrong evidence.
Are repeats used appropriately?Repeated trials can reveal consistency, but do not fix every design problem.
Are enough suitable specimens used?Claims about a wider group need evidence that captures natural variation.
Is the tested range suitable?A trend claim needs more than one narrow comparison.
Is there a useful reference or control?A reference can help isolate what changed or rule out alternatives.
Does the conclusion stay within the evidence?Strong measurements can still be weakened by an overbroad claim.

Worked Example 1 — One Fair Trial Versus Ten Unfair Trials

Claim: A larger exposed wet surface causes faster water loss by evaporation under otherwise similar conditions.

Investigation A: Two identical wet cloths contain equal starting amounts of water, remain in the same surroundings for the same time, but one is spread out and one is folded. Each condition is tested once.

Investigation B: One cloth is spread out in moving air near a window; the other is folded in still air elsewhere. Each condition is repeated ten times.

Which gives stronger evidence about surface area?

Investigation A has the stronger causal comparison because surface area is better isolated. Investigation B has more repeated data, but air movement is another difference that could affect water loss.

Repeating a confounded comparison gives more measurements of a weak comparison.

Best design: keep the fair comparison of A, then repeat it appropriately.

Worked Example 2 — One Specimen Repeated Many Times Versus Several Similar Specimens

Claim: Leaves of a particular kind generally show a certain response under a condition.

Investigation A: The same single leaf is tested repeatedly.

Investigation B: Several similar leaves are tested under the same planned condition.

If the claim is about leaves of that kind generally, Investigation B may give stronger evidence because it samples biological variation. Repeating one leaf can tell you more about consistency for that leaf, but not necessarily about the wider group.

If the claim were instead about the repeatability of one sensor reading on that same leaf, repeated measurements could serve a different evidence job.

Worked Example 3 — Endpoint Measurement Versus Time Series

Claim: Set-up P changes faster than Set-up Q during the first ten minutes.

Investigation A: Measures only the final value after thirty minutes.

Investigation B: Measures both set-ups at 0, 5 and 10 minutes using the same method.

For the first-ten-minutes rate claim, B gives stronger evidence because the measurements align with the interval named in the claim.

For a different claim—“Which has the higher final value after thirty minutes?”—A may be sufficient.

Worked Example 4 — Direct Measurement Versus an Indirect Indicator

Claim: A process occurs more under Condition X than Condition Y.

Investigation A: Measures a quantity directly related to the process in this setup.

Investigation B: Observes a colour change that is only an indirect indicator and can also be influenced by another condition.

A may provide stronger evidence if its measurement is valid, comparable and directly tied to the claim.

But “direct” is not automatically better. A poorly measured direct quantity can be weaker than a reliable validated indicator. The learner must ask what each measure actually supports.

Worked Example 5 — Two Conditions Versus a Wider Range

Claim: As light level increases, the measured response generally increases across the tested range.

Investigation A: Tests only a low and a high light condition.

Investigation B: Tests five ordered light levels with comparable methods.

B gives stronger evidence for a trend across a range because it shows what happens between the endpoints. A can establish a difference between two tested conditions but cannot reveal whether the relationship is smooth, has a plateau, reverses or contains an exception.

Worked Example 6 — Reference Set-Up Versus No Reference

Claim: A treatment changes the measured outcome compared with the ordinary condition.

Investigation A: Measures only the treated set-up.

Investigation B: Measures the treated set-up and an otherwise comparable reference set-up without the treatment.

B may provide stronger evidence because the reference helps show what would happen without the tested treatment under comparable conditions.

A control is not magical and is not required in exactly the same form for every investigation. Its value depends on what alternative explanation or baseline it helps test.

Worked Example 7 — Precise Instrument, Wrong Outcome

Claim: A surface affects how far a toy car travels.

Investigation A: Uses a highly precise thermometer to measure air temperature but does not measure distance travelled.

Investigation B: Uses a simple ruler appropriately to measure the car’s travel distance while controlling relevant starting conditions.

B gives stronger evidence for the travel-distance claim. Measurement precision matters only after the measured quantity is relevant.

High precision in the wrong quantity is not strong evidence for the claim.

Worked Example 8 — More Data Points but a Changed Measuring Method

Investigation A records ten time points, but every measurement requires opening a sealed container and disturbing its contents. Investigation B records fewer time points using a method that leaves the condition intact.

If opening the container can change the process being investigated, B may provide stronger evidence despite fewer measurements.

Evidence quantity and evidence quality are not the same thing.

Worked Example 9 — Strong Data, Overbroad Conclusion

An investigation carefully compares two materials under controlled conditions and finds Material P performs better than Material Q in that test.

The student concludes: “Material P is the best material for all uses.”

The measurements may be strong for the tested comparison. The conclusion is weak because it travels beyond the tested materials, conditions and performance measure.

Strong evidence does not authorise an unlimited conclusion.

Worked Example 10 — Same Claim, Different Evidence Jobs

Claim: A changed condition affects a process.

One investigation may be stronger for causal attribution because it controls alternative conditions. Another may be stronger for generalisation because it tests more varied specimens. A third may be stronger for time pattern because it records repeated measurements.

There may not be one investigation that is strongest in every dimension. The learner should identify which evidence property matters most for the claim being judged.

There Is No Universal Evidence Ladder

Do not memorise a fixed hierarchy such as:

more repeats > control > more specimens > more conditions.

That is not scientifically defensible because each feature solves a different problem.

Evidence weaknessPossible improvement
Trial-to-trial variationRepeat the trial appropriately
Natural specimen variationUse more suitable similar specimens
Cannot see a trendTest more ordered condition values
Cannot isolate a causeImprove fair-comparison control
No baseline or referenceAdd a scientifically useful comparison set-up
Measured the wrong outcomeChange the measurement to match the claim
Cannot resolve timingUse suitable repeated measurements over time
Measurement disturbs the systemUse a less intrusive method

Claim–Method Alignment Comes First

Before checking repeats or precision, ask whether the investigation can answer the claim at all.

Examples:

  • A final measurement cannot establish the exact time an event began.
  • One categorical comparison cannot establish a continuous numerical trend.
  • A comparison with two changed conditions cannot isolate one cause.
  • One specimen cannot strongly represent variation across a diverse group.
  • A measurement of temperature does not automatically answer a claim about amount of material.

If claim–method alignment fails, additional polish cannot repair the central problem.

Fairness Before Repetition

Repeated trials are valuable when the underlying comparison is meaningful.

If two set-ups differ in several relevant conditions, repeating the same confounded comparison does not isolate one factor. The learner should first repair the comparison, then use repetition if variation or consistency is still an evidence problem.

Relevance Before Precision

A measurement can have many decimal places and still be irrelevant to the claim.

Ask:

Does this measured quantity actually change if the claim is true?

If not, the instrument’s precision does not make the evidence stronger for that claim.

Range Before Trend

Two points can show a difference. Several ordered conditions are usually more informative when the claim concerns a trend, threshold, plateau or turning point.

More conditions do not automatically make a study better either. The range must be scientifically meaningful and measured consistently.

Specimens Before Generalisation

If the conclusion is about many organisms, leaves, seeds or natural samples, evidence from only one specimen may be fragile because natural variation exists.

But if the question is about one specific object, adding many different specimens may answer a different question.

Control or Reference: Ask What It Does

Do not write “Investigation B is stronger because it has a control” without explaining the control’s scientific job.

A useful control/reference may:

  • provide a baseline;
  • show what happens without the tested condition;
  • help rule out an alternative explanation;
  • confirm that another part of the procedure behaves as expected.

The strength comes from the comparison it enables.

More Measurements Are Not the Same as More Independent Evidence

Ten measurements of the same ongoing trial provide a detailed time series. Ten independently repeated trials provide information about repeatability. Ten specimens provide information about variation across specimens.

These are different evidence jobs.

Count what the data represent, not merely how many rows are in the table.

The Stronger-Evidence Comparison Protocol

  1. Write the exact claim.
  2. For Investigation A, name the changed condition.
  3. Name its measured outcome.
  4. Repeat for Investigation B.
  5. Check fair-comparison quality.
  6. Check whether each outcome is relevant to the claim.
  7. Check repeats and consistency.
  8. Check specimen choice and variation.
  9. Check the range of tested conditions.
  10. Check control/reference value.
  11. Check measurement resolution and interference.
  12. Check whether each conclusion stays within what was tested.
  13. Choose the stronger evidence for the stated claim.
  14. Explain the decisive reason, not just the number of features.

The Decisive-Difference Rule

When two investigations differ in many ways, do not list every difference equally. Identify the one or two differences that most affect whether the claim is supported.

Example:

“Investigation B gives stronger evidence for the effect of surface area because only surface area differs between the comparison set-ups, whereas Investigation A also changes air movement. B therefore better isolates the condition named in the claim.”

This is stronger than: “B is better because it is fairer.”

Observable Failure Signatures

Failure signatureEarliest weak linkRepair
“B has more readings, so B is stronger.”Data quantity substituted for claim alignment.Check what the readings actually measure.
“A repeats ten times, so it is fair.”Repeatability confused with variable control.Inspect which conditions differ.
“This instrument is more advanced, so the evidence is better.”Apparatus complexity confused with relevance.Ask whether the quantity measured answers the claim.
“One leaf was tested carefully, so the result applies to all leaves.”Measurement quality confused with representativeness.Check specimen variation and conclusion scope.
“The control means nothing happens there.”Control purpose misunderstood.State what comparison the control enables.
“Five conditions are always better than two.”Range treated as a universal quality rule.Match number/range of conditions to the claim.
“A has smaller error bars / variation, therefore its causal claim is stronger.”Consistency substituted for causal isolation.Check fairness and relevance before variation.

The Earliest Weak-Link Diagnosis

  • Claim: Did I state exactly what is being tested?
  • Changed condition: Does the design isolate it?
  • Measured outcome: Does it answer the claim?
  • Comparison: Are relevant conditions controlled?
  • Repeatability: Do repeats address variation rather than hide a design flaw?
  • Specimens: Does the sample match the conclusion scope?
  • Range: Is there enough condition coverage for the claimed pattern?
  • Reference: Does a control/baseline rule out an important alternative?
  • Measurement: Is it suitable and non-interfering?
  • Conclusion: Does it stay within what was actually tested?

Misconception Repair — “More Data Always Means Stronger Evidence”

No. More relevant, well-collected data can strengthen evidence. More irrelevant or confounded data can simply make a weak design larger.

Misconception Repair — “Repeating Fixes an Unfair Test”

No. Repeating an unfair comparison does not make the changed factors easier to separate. Repair the design first.

Misconception Repair — “A Control Is Always Required”

Not every scientific question uses a separate control set-up in the same way. The important question is whether the comparison provides the baseline or alternative needed to test the claim.

Misconception Repair — “The Most Precise Instrument Gives the Best Investigation”

Precision is useful only when the instrument measures the right quantity, has a suitable range and does not disturb the system in a way that changes the outcome.

Misconception Repair — “Strong Evidence Proves the Claim Forever”

Evidence supports conclusions within tested conditions and reasonable scientific limits. One well-designed Primary Science investigation does not establish a universal law across every organism, material or environment.

Misconception Repair — “A Bigger Difference Means Stronger Evidence”

A large observed difference can still come from an unfair comparison. Evidence strength depends on design and measurement, not only effect size.

The Claim–Evidence Match Table

If the claim is about…Stronger evidence usually needs…
A causal effect of one conditionA fair comparison that isolates that condition
Consistency of a resultAppropriate repeated trials or measurements
A wider group of natural specimensSeveral suitable specimens representing relevant variation
A trend across a conditionSeveral ordered condition values across a meaningful range
Timing or rate changeMeasurements at suitable time intervals
A difference from ordinary/baseline conditionA meaningful reference or control comparison
A precise quantityA suitable instrument with adequate scale/resolution
A process inferred from an indicatorEvidence that the indicator validly tracks that process under the stated conditions

Question-Reading Protocol for “Which Investigation Is Better?”

  1. Read the scientific claim before reading the methods.
  2. Underline the condition the claim is about.
  3. Circle the outcome that should respond if the claim is true.
  4. Read Investigation A and ask whether it changes/measures those exact things.
  5. Read Investigation B the same way.
  6. Mark any extra changed conditions.
  7. Mark any missing baseline, range or repetition that matters.
  8. Identify one decisive advantage.
  9. Explain how that advantage improves the evidence for the claim.
  10. State the limit that remains.

Practice Sequence

  1. Claim matching: pair claims with suitable measured outcomes.
  2. Fairness comparison: choose between a fair small study and an unfair large study.
  3. Repeat-versus-specimen practice: identify which kind of variation matters.
  4. Range practice: decide when two conditions are enough and when a trend needs more.
  5. Control reasoning: explain what a reference set-up contributes.
  6. Measurement relevance: reject precise but irrelevant measurements.
  7. Conclusion-scope practice: narrow overgeneralised claims.
  8. Mixed comparison: rank two investigations and state the decisive reason.
  9. Transfer: repeat in unfamiliar scientific contexts.

Unfamiliar Transfer Challenge

A mystery material is tested to see whether Condition X affects how quickly its mass changes.

Investigation A: Tests X and no-X under otherwise similar conditions. Measures mass at the start and after 20 minutes. Repeats three times.

Investigation B: Tests six different levels of X, but each level is kept at a different temperature. Measures final mass only once.

Which gives stronger evidence that X itself affects mass change?

A is stronger for isolating the causal effect of X because temperature is controlled and the start-to-finish change is measured comparably. B has a wider condition range, but temperature is confounded with X.

Could B become useful? Yes. If temperature were controlled, its wider range could provide stronger evidence about the shape of the relationship across X.

Delayed Independent Return

Three to five days later, compare two fresh investigations without notes. Write a short evidence receipt:

  • Exact claim:
  • Investigation A — changed condition:
  • Investigation A — measured outcome:
  • Investigation B — changed condition:
  • Investigation B — measured outcome:
  • Which comparison is fairer?
  • Which measurement is more relevant?
  • Which better addresses repeatability or variation?
  • Which range/specimen choice better matches the claim?
  • Strongest investigation for this claim:
  • Decisive reason:
  • Remaining limitation:

The Stronger-Evidence Receipt

  • I stated the claim before ranking the investigations.
  • I identified the changed condition.
  • I identified the measured outcome.
  • I checked fair-comparison logic.
  • I checked measurement relevance before precision.
  • I separated repeated trials from repeated time points and more specimens.
  • I checked whether the tested range matches the claim.
  • I explained the job of any control/reference set-up.
  • I checked whether measurement changes the system.
  • I checked conclusion scope and generalisation.
  • I selected the stronger evidence for the specific claim—not the “fancier” experiment.
  • I named the decisive reason.

Evidence and Model Limits

Real scientific evidence evaluation can involve sample size calculations, uncertainty estimates, statistical analysis, replication across laboratories, calibration, randomisation and many other methods. Primary learners do not need that machinery.

The age-appropriate principle is still powerful: evidence strength depends on how well the investigation links a scientific claim to a fair, relevant and interpretable observation or measurement.

There is also no guarantee that one investigation is stronger in every respect. A learner should identify which design features matter for the claim and state any remaining limitations.

Useful Internal Routes

Parent and Tutor Teaching Guide

When a learner says, “Investigation B is better,” do not accept the label by itself. Ask:

“Better evidence for which exact claim?”

Then ask for one decisive reason. Useful responses sound like:

  • “B isolates the tested condition because the other relevant conditions are kept comparable.”
  • “A measures the outcome that directly answers the claim.”
  • “B uses several specimens, so it better supports a conclusion about that type of organism rather than one individual.”
  • “A tests more ordered conditions, so it gives stronger evidence for a trend rather than one pairwise difference.”

Create paired investigations where the “bigger” experiment is intentionally weaker. This prevents children from learning superficial rules such as “more readings = better”.

Then reverse the task. Give one strong investigation and ask what kind of claim it can support. This teaches that evidence strength and conclusion scope must be matched in both directions.

Finally, use delayed transfer. Several days later, present an unfamiliar setup and ask the learner to rank the evidence without any hints about repeats, controls or specimens. The learner should diagnose the claim first.

Authoritative and Research References

The Quiet Ending

Strong evidence is not the experiment with the most equipment.

It is not automatically the one with the most numbers.

It is the investigation whose design lets the observations speak clearly to the claim.

Start with the claim. Follow the method. Check the evidence. Then decide how far the conclusion deserves to travel.