Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

How to Perform in the new G2 SEC Examinations | Learner’s Guide Vol 0036 | Science: Evidence Strength — Observation, Inference, Causation and Conclusions

How to perform in the new G2 SEC Science examination at an advanced level requires the learner to control not only scientific knowledge but the strength of the claim being made. An observation is not the same as an inference. A correlation is not automatically a cause. A plausible mechanism is not automatically proven by one data point. A conclusion should be no broader than the evidence.

This thirty-sixth Learner’s Guide focuses on evidence strength: observation, measurement, pattern, inference, mechanism, causation and conclusion. The core rule is: say exactly what the evidence supports—no less, no more.

For 2027, SEAB lists G2 Science as K223 Science (Physics, Chemistry), K224 Science (Physics, Biology) and K225 Science (Chemistry, Biology). The shared assessment objectives include interpreting information, identifying patterns, drawing inferences, providing reasoned explanations, making predictions and evaluating data and methods. Use the SEAB 2027 G2 syllabus page for current official details.

The Evidence Ladder

Scientific answers operate at different levels. Confusion occurs when the learner jumps levels without justification.

  • Observation: what was seen, measured or recorded.
  • Pattern: how observations relate across values or conditions.
  • Inference: a conclusion supported by the pattern and scientific knowledge.
  • Mechanism: the scientific process explaining why the pattern occurs.
  • Causal claim: a statement that changing one factor produces a change in another.
  • Conclusion: the final claim supported by the investigation or evidence.

Each step requires enough evidence to justify moving upward.

Observation: Stay Close to What Was Detected

An observation is direct: colour changed, temperature increased, a precipitate formed, the pointer moved, the plant grew taller, the graph value was larger.

Do not insert a hidden explanation into an observation answer. “Carbon dioxide was produced” may be an inference if the actual observation was that limewater turned milky after gas was passed through it.

Measurement Is Quantified Observation

A measurement gives a number with a quantity and unit. It is stronger than a vague description because it makes comparison possible.

However, measurement quality depends on apparatus, resolution, method and repeatability. A number is not automatically perfect evidence.

Pattern: More Than One Point

A pattern describes how values behave across conditions: increase, decrease, plateau, optimum, direct trend, inverse trend or another relationship.

Do not claim a trend from one isolated comparison if the question provides a larger data set. Use the full pattern.

Inference: The Smallest Justified Step

An inference goes beyond direct observation but remains anchored to evidence and known Science.

A strong inference is conservative: it says what the evidence suggests, not everything that could possibly be true.

Mechanism: Why the Pattern Happens

A mechanism links the evidence to scientific process: collisions, force relationships, energy transfer, diffusion, enzyme behaviour, electrical effects or another syllabus model.

The mechanism should explain the observed relationship rather than replace it. State the evidence first when the question provides data.

Causation: A Stronger Claim

To claim that X causes Y, the design must support more than simple co-occurrence. The learner should consider whether X was changed, whether relevant alternatives were controlled and whether the observed change in Y can reasonably be attributed to X.

A graph where X and Y rise together does not automatically prove causation.

Correlation

Correlation means variables are associated. Positive correlation means they tend to increase together; negative correlation means one tends to decrease as the other increases.

Correlation can be scientifically useful, but it does not by itself identify the mechanism or rule out other variables.

The Causation Check

  • Was the proposed cause deliberately changed?
  • Was the outcome measured?
  • Were important alternative variables controlled?
  • Is there a plausible scientific mechanism?
  • Are repeats or multiple data points available?
  • Does the conclusion stay within the tested conditions?

The more of these conditions are satisfied, the stronger the causal interpretation can be.

Conclusion: Answer the Aim

A conclusion should return to the investigation question or aim. It should identify the relationship the data supports and avoid claims beyond the evidence.

A conclusion is not simply the final data point.

The Strength-of-Claim Vocabulary

Language can reflect evidence strength:

  • shows — use when the evidence directly demonstrates the stated relationship within the tested context;
  • supports — useful when evidence is consistent with a conclusion;
  • suggests — appropriate when evidence points toward an explanation but remains limited;
  • is consistent with — useful when several explanations may still be possible;
  • does not establish — useful when evidence is insufficient for a stronger claim.

The learner should not weaken every answer with “maybe”, but should match language to the evidence.

One Measurement Is Weak Evidence

A single reading may be affected by random variation or measurement error. Repeated measurements can provide stronger evidence about consistency.

This does not mean a single reading is useless. It means the conclusion should reflect its limited reliability.

Repeated Measurements

Repeats can reveal variation, identify anomalies and support a mean where appropriate.

Repeats strengthen reliability but do not automatically remove systematic error. If the instrument is biased, every repeat may be biased in the same direction.

Sample Size

In biological or population contexts, a larger appropriate sample can reduce the influence of individual variation and provide stronger evidence about a broader group.

The required sample depends on the question. Do not use “larger sample size” as a generic phrase without explaining why variation matters.

Range of the Independent Variable

A wider useful range can reveal whether a relationship holds across more conditions and may expose turning points or plateaus.

A narrow range supports conclusions only within that narrow range.

Number of Data Points

More levels of the independent variable can reveal the shape of a relationship. This is different from taking more repeats at the same level.

Both can strengthen evidence, but in different ways.

Anomalies

An anomaly is evidence too. Do not delete it automatically. Ask whether it resulted from procedural error, random variation or a real feature of the system.

A repeat at the same condition can help determine whether the unusual result is reproducible.

Systematic Bias

Systematic bias weakens accuracy even if repeated readings are consistent. Examples can include zero error, heat loss or consistent material loss depending on the experiment.

Strong evaluation identifies how the bias affects the result and how the method could be improved.

Evidence Quality and Apparatus

Range, resolution and suitability of apparatus affect evidence strength. A coarse instrument may not distinguish small differences; an instrument with insufficient range may fail entirely.

Measurement quality belongs in the conclusion only when it materially affects confidence in the result.

Evidence Quality and Controls

Controlled variables strengthen interpretation by reducing alternative explanations.

A control is scientifically useful when the variable could otherwise influence the dependent outcome.

Evidence Quality and Fair Comparison

A comparison is fair when relevant conditions are sufficiently similar except for the variable being investigated.

“Fair test” should be translated into specific controlled conditions rather than used as a slogan.

Physics Evidence Strength

Physics often uses quantitative evidence, which can create a false sense of certainty. A precise number is still only as strong as the measurement, model and assumptions behind it.

A speed calculated from measured distance and time may be reliable within the measurement conditions, but claiming the object will continue at that speed indefinitely requires an additional assumption.

Physics: Model Versus Reality

Physical models simplify. Friction may be ignored, rates may be treated as constant and objects may be idealised. The learner should know whether the question asks for the model result or evaluation of the real situation.

Evidence strength changes when model assumptions are relaxed.

Physics: Graph Evidence

A straight-line graph can support a linear relationship within the measured range. The gradient can quantify that relationship where appropriate.

Do not assume the same line continues outside the measured range without justification.

Chemistry Evidence Strength

Chemistry often uses observations as evidence of substances or reactions. A colour change, precipitate, gas or temperature change may support an inference, but the learner should know whether the observation is unique to one explanation.

Where several causes could produce the same observation, stronger identification may require another test or piece of evidence.

Chemistry: Observation Versus Chemical Conclusion

“A white precipitate formed” is an observation. “A particular ion is present” is an inference based on the test conditions and known chemistry.

Do not collapse the two when the command word asks specifically for observation.

Chemistry: Reaction Evidence

A temperature increase may support the conclusion that energy was released to the surroundings, but the full chemical interpretation depends on the system and reaction context.

Use the evidence given rather than writing a memorised reaction story that ignores the measured result.

Biology Evidence Strength

Biology frequently contains natural variation. One organism, one leaf or one trial may not represent the entire population.

Conclusions should consider sample size, biological differences, environmental control and whether repeated patterns are present.

Biology: Variation Is Not Automatically Error

If repeated biological measurements differ, the difference may reflect genuine variation rather than careless measurement.

The learner should distinguish measurement uncertainty from natural biological diversity.

Biology: Association and Cause

A study may show that two biological variables are associated. Unless the design manipulates the proposed cause and controls alternatives, the evidence may support correlation rather than causation.

This distinction is especially important in data-response questions based on observational studies.

The Evidence Triangle

  • Quantity: how much evidence is available?
  • Quality: how reliable and accurate is the evidence?
  • Relevance: does the evidence actually answer the claim?

A large amount of irrelevant data does not strengthen the conclusion. One highly relevant measurement may be more useful than many unrelated observations.

Quantity of Evidence

Quantity includes number of measurements, repeats, data points or samples. More evidence can strengthen confidence when collected appropriately.

However, repeating a flawed method many times does not fix systematic bias.

Quality of Evidence

Quality includes measurement resolution, control of variables, repeatability, method suitability and whether the evidence was gathered consistently.

A strong conclusion depends on the quality of the process producing the data.

Relevance of Evidence

Evidence must connect directly to the claim. A student may cite a true scientific fact that does not support the specific conclusion.

Ask: if this evidence were removed, would the claim become weaker? If not, it may be irrelevant.

The Alternative-Explanation Test

Before making a strong causal claim, ask whether another variable or mechanism could explain the same pattern.

If a reasonable alternative remains uncontrolled, use more cautious language or identify what further evidence would be needed.

The Falsification Question

Ask what evidence would show the conclusion is wrong. This is a powerful way to test whether the claim is genuinely linked to the data.

If no imaginable result could change the conclusion, the reasoning may be too vague or unfalsifiable for the examination task.

The Replication Question

Would repeating the investigation under the same conditions likely produce a similar pattern? If not, the evidence may be too unstable for a strong conclusion.

Replication is especially useful for evaluating reliability.

The Generalisation Question

Can the conclusion apply only to the tested samples and conditions, or to a wider group?

Wider generalisation requires broader evidence. Do not move from one narrow experiment to a universal statement without support.

The Mechanism Question

Does a plausible syllabus-level mechanism connect cause and outcome? Mechanism alone does not prove causation, but it strengthens the scientific coherence of the explanation.

A mechanism unsupported by the data should not replace the data.

Claim Strength in Data Questions

When asked to conclude from a graph, begin with the pattern the graph directly supports. Then move to explanation only if the command asks for it.

This preserves the distinction between evidence and interpretation.

Claim Strength in Experimental Questions

When asked whether a method supports a conclusion, inspect variables, repeats, range, anomalies and measurement quality.

A conclusion may be reasonable but still weakly supported if the method cannot separate alternative explanations.

Claim Strength in Unfamiliar Contexts

Unfamiliar contexts tempt learners to rely on general knowledge. Instead, use the question’s evidence first and syllabus principles second.

The answer should not become broader simply because the learner knows extra facts.

The Strong-Claim Warning Words

  • always;
  • never;
  • proves;
  • definitely;
  • all;
  • none;
  • causes.

These words may be correct in some contexts, but they require strong support. During checking, inspect any absolute claim carefully.

The Weak-Claim Warning

The opposite problem is excessive caution: “maybe”, “perhaps” and “might” in every sentence.

If the data clearly supports a pattern within the tested conditions, state it directly. Scientific caution should be precise, not timid.

The Evidence-to-Claim Match

A useful rule is to make the claim only one level stronger than the evidence warrants, never several levels stronger.

Observation can support a description. Repeated pattern can support an inference. Controlled manipulation with a plausible mechanism can support a stronger causal interpretation.

The Evidence Audit

  1. What is directly observed or measured?
  2. What pattern exists?
  3. What inference is justified?
  4. What mechanism explains it?
  5. What alternative explanations remain?
  6. How broad can the final conclusion be?

This six-step audit is useful in longer structured questions.

The One-Claim Drill

Give the learner a data set and require exactly one conclusion. The conclusion must be as strong as possible without exceeding the evidence.

Then ask what extra evidence would be needed for a stronger claim.

The Observation-Inference Drill

Provide ten statements and classify each as observation or inference. Then identify what evidence supports each inference.

This is especially useful for Chemistry and experimental Biology.

The Correlation-Causation Drill

Use graphs showing associated variables. Ask whether the evidence establishes causation, and what additional design features would strengthen a causal claim.

The learner practises restraint and experimental reasoning together.

The Conclusion-Boundary Drill

Take broad conclusions and rewrite them so they match the tested range, sample or conditions.

This trains language such as “within the tested range” and “for the samples used” where appropriate.

The Alternative-Explanation Drill

Give an experiment with one uncontrolled variable. Ask the learner to identify how that variable could provide another explanation for the result.

Then suggest the relevant control or redesign.

The Evidence-Quality Ranking Drill

Present three studies or data sets with different sample sizes, repeats, ranges and measurement quality. Rank which provides stronger evidence and explain why.

This turns evaluation from generic criticism into comparative judgement.

The Claim-Language Drill

Rewrite the same conclusion using “shows”, “supports”, “suggests” and “is consistent with”. Discuss which wording best matches the evidence.

The learner develops control over scientific certainty.

The Evidence Error Ledger

  • observation written as inference;
  • one point treated as a trend;
  • correlation treated as causation;
  • mechanism stated without evidence;
  • conclusion broader than tested range;
  • repeat used to claim removal of systematic error;
  • sample variation ignored;
  • absolute wording stronger than evidence;
  • relevant limitation identified but consequence not explained;
  • general knowledge used instead of question evidence.

These errors should be reviewed separately from content recall.

A Four-Week Evidence-Strength Build

Week 1 — observation and pattern

Separate direct evidence from interpretation and practise accurate descriptions.

Week 2 — inference and mechanism

Build evidence → concept → mechanism chains without overclaiming.

Week 3 — causation and evaluation

Use experiments, controls, repeats, sample size and alternative explanations.

Week 4 — timed structured transfer

Use mixed K223/K224/K225-style structured questions and check claim strength under time.

Use Structured-Question Decomposition

Use Vol 0032. Evidence strength becomes easier to control when task, evidence, operation and answer are separated.

Use Experimental-Question Control

Use Vol 0020 for variables, measurement quality, reliability and method improvements.

Use Command-Word Control

Use Vol 0024. The command word determines whether the evidence should be described, explained, predicted from, justified or evaluated.

Use Examination Craft

For pacing and checking, continue through the Examination Craft hub. Evidence control must remain precise even late in the paper.

The PSLE Bridge

The PSLE rule Evidence Before Explanation remains the foundation. At G2, the learner adds a second discipline: match the strength of the explanation and conclusion to the strength of the evidence.

Final Rule

Do not make the claim stronger than the evidence.

Observe accurately. Describe patterns honestly. Infer cautiously. Explain with the correct mechanism. Distinguish association from cause. Evaluate evidence quality. Conclude within the tested boundary. Scientific performance is not only knowing what might be true—it is knowing what the evidence allows you to say.

Scenario 1 — One Point on a Graph

A graph contains one high value at x = 5. The learner writes, “Y increases as X increases.” One point cannot establish a trend by itself. The learner should inspect the rest of the data before making the pattern claim.

If only one comparison is available, describe that comparison rather than inventing a broader relationship.

Scenario 2 — A Repeated Pattern

Several points show Y increasing as X increases, with small variation. The evidence now supports a trend statement within the measured range.

If the question asks why, the learner can add the relevant mechanism. The mechanism should explain the trend rather than replace it.

Scenario 3 — Correlation in Biology

A data set shows students who sleep more tend to perform better on a task. This supports an association. It does not prove sleep alone caused the performance difference because other variables may differ.

A stronger causal claim would require a design that addresses alternative explanations and appropriate ethical constraints.

Scenario 4 — Controlled Physics Experiment

A Physics experiment deliberately changes one quantity, keeps relevant conditions constant and records a consistent change in another quantity across several levels. The design provides stronger evidence for a causal relationship within the tested model.

The conclusion should still stay inside the tested range and assumptions.

Scenario 5 — Chemistry Identification

A gas turns limewater milky. The direct observation is the change in limewater. The chemical inference is that carbon dioxide may be present under the known test conditions.

If another independent test supports the same identification, the evidence becomes stronger.

Scenario 6 — Anomalous Result

Five repeated measurements cluster closely but one is very different. The learner should not average blindly without considering the anomaly.

Check whether a procedural reason exists, repeat the condition if possible and decide whether the unusual value reflects error or real variation.

Scenario 7 — Systematic Bias

Every temperature measurement is taken with an instrument that reads 2°C too high. The readings may be precise relative to one another but inaccurate.

Repeating the experiment with the same biased instrument does not remove the systematic error.

Scenario 8 — Narrow Range

An experiment tests a relationship only between 20 and 30 units. The learner concludes the same trend occurs from 0 to 100.

The conclusion is too broad. Evidence supports only the measured interval unless a validated model justifies wider extension.

Scenario 9 — Small Biological Sample

Two leaves from one plant show a pattern. The learner concludes all plants of the species behave the same way.

The sample is too limited for that broad generalisation. The answer should stay close to the tested samples or identify the need for broader sampling.

Scenario 10 — Mechanism Without Data

The learner knows a strong theoretical mechanism and writes it even though the supplied graph does not show the expected pattern.

Theory should not override the evidence provided. First describe what the graph actually shows, then discuss why the result may differ if the question asks for evaluation.

Scenario 11 — Data Without Mechanism

The learner accurately describes the pattern but the command word asks “explain”. The answer is incomplete because evidence has not been connected to scientific cause.

Add the mechanism at the correct syllabus level.

Scenario 12 — Strong Wording From Weak Evidence

A single uncontrolled comparison is described as “proving” a cause. Replace the wording with a more defensible claim such as “suggests an association” if that is what the evidence supports.

Claim language should reflect design strength.

Scenario 13 — Excessive Caution

A well-controlled experiment produces a clear consistent relationship, but the learner writes “maybe there could perhaps be a slight relationship”.

This understates the evidence. Scientific caution should not prevent direct conclusions where the evidence is strong.

Scenario 14 — Prediction Beyond the Data

A graph rises steadily across the measured range. The learner predicts that it will rise forever. The prediction exceeds the evidence.

A better answer limits the prediction to a reasonable nearby extension if the question and model support extrapolation.

Scenario 15 — Method Improvement and Claim Strength

A method is improved by controlling temperature and increasing repeats. The learner should understand why the conclusion can now be stronger: fewer alternative explanations and better evidence about consistency.

Method improvement is valuable because it changes the evidence quality, not because “more accurate” is a magic phrase.

The Claim-Strength Scale

  • Level 1 — direct observation;
  • Level 2 — pattern within data;
  • Level 3 — evidence-based inference;
  • Level 4 — mechanism-supported explanation;
  • Level 5 — causal claim supported by appropriate design;
  • Level 6 — broader conclusion supported by sufficient range, sample and replication.

Not every question requires the highest level. Use the level the evidence can support.

The Evidence Sufficiency Question

Ask: What evidence would I need before making this claim stronger?

The answer may be more repeats, a wider range, better controls, another independent test, a larger sample or a measurement with suitable resolution.

The Competing-Explanation Question

Ask: What else could produce this result?

If a plausible alternative remains, the conclusion should be cautious or the experiment should be redesigned to distinguish the explanations.

The Direction-of-Error Question

When a limitation is identified, ask how it would affect the result. Would the value be too high, too low, more variable or simply uncertain?

Explaining direction makes evaluation more scientific and specific.

The Consistency Question

Do repeated measurements tell the same story? Consistency strengthens confidence, but the learner should still consider systematic bias and shared method flaws.

The Independent-Evidence Question

Does another measurement, test or representation support the same conclusion? Independent evidence can strengthen confidence because it does not rely on exactly the same failure mode.

This mirrors the checking principle used in Mathematics.

The Scope Question

Exactly what population, range, system or set of conditions does the conclusion apply to?

Writing the scope explicitly prevents accidental overgeneralisation.

The Evidence Table

For complex questions, use a small mental or written table:

  • claim;
  • supporting evidence;
  • limitation;
  • appropriate wording.

This is especially useful in evaluation questions.

Evidence Strength in MCQ

Multiple-choice questions may present several true statements. Choose the option best supported by the evidence in the stem, not the option containing the most impressive science vocabulary.

Eliminate options that claim more than the data allows.

Evidence Strength in Structured Questions

Structured responses make reasoning visible. State the relevant evidence, then the inference or mechanism. Do not make the examiner infer the connection for you.

Evidence Strength Under Time

Time pressure encourages shortcut claims. Use a compact rule: data first, claim second.

If the claim contains an absolute word or causal statement, perform one quick evidence check before moving on.

The Evidence Check in the Final Pass

  • Did I confuse observation with inference?
  • Did I claim a trend from too little data?
  • Did I claim causation from correlation alone?
  • Did I ignore an anomaly or major limitation?
  • Did I generalise beyond the tested range or sample?
  • Is my certainty language stronger or weaker than the evidence?

These checks are high value because they target reasoning, not cosmetic wording.

The Evidence-Strength Error Ledger

  • observation replaced by interpretation;
  • pattern claimed from insufficient points;
  • correlation treated as cause;
  • mechanism stated without linking evidence;
  • systematic error treated as random;
  • repeats claimed to improve accuracy automatically;
  • sample too small for conclusion;
  • range too narrow for generalisation;
  • absolute wording unsupported;
  • strong evidence weakened unnecessarily by vague language.

The ledger turns “Science reasoning weak” into specific repair tasks.

The 15-Minute Evidence Session

  1. Five minutes: classify observation, pattern, inference and mechanism in sample statements.
  2. Five minutes: review one graph or table and write the strongest justified conclusion.
  3. Five minutes: identify one limitation and state how it changes confidence in the conclusion.

This develops evidence judgement without requiring a full paper.

The Experimental-Evidence Session

Take one method and ask what design feature supports causation, what threatens reliability, what threatens accuracy and what further evidence would strengthen the conclusion.

This integrates variables, measurement quality and claim strength.

The Cross-Discipline Evidence Session

Use one Physics, one Chemistry and one Biology example. Compare what counts as strong evidence in each discipline.

The shared principle is the same: claim strength must match evidence quality, even though the measurements and mechanisms differ.

A Four-Week Evidence-Strength Build

Week 1: observation and pattern. Week 2: inference and mechanism. Week 3: causation, reliability and alternative explanations. Week 4: timed mixed structured questions with claim-language checking.

Use the Science Estate for Concept Repair

If the learner cannot judge evidence because the underlying concept is missing, return to the Complete Science Index. Evidence reasoning cannot substitute for missing scientific knowledge.

Use Examination Craft

For paper-level pacing and checking, continue through the Examination Craft hub. The evidence check should become fast enough to survive the final part of the paper.

Final Standard

The advanced G2 Science learner can distinguish what was observed, what was inferred, what mechanism explains it, what alternative explanations remain and how broad the final claim can be.

They do not confuse confidence with evidence, and they do not weaken strong data unnecessarily. The final answer says exactly what the Science allows.

The Final Evidence Calibration Checklist

  • What was directly observed or measured?
  • What pattern is actually present?
  • What inference follows without adding unsupported assumptions?
  • What mechanism explains the pattern?
  • Does the design justify a causal claim or only an association?
  • What limitation most affects confidence?
  • How far can the conclusion reasonably generalise?
  • Does the wording match the strength of the evidence?

This checklist is most useful for longer data, experiment and evaluation questions. The learner should not write every step explicitly; the aim is to make the judgement automatic.

The Claim-Strength Finish

When checking a conclusion, replace the question “Does this sound scientific?” with “What evidence on this page supports this exact claim?” If the answer is weak, narrow the claim or identify the missing evidence. If the evidence is strong, state the conclusion directly rather than hiding behind unnecessary uncertainty.

Scientific maturity is calibrated certainty: neither overclaiming nor underclaiming.

Final Perspective

Evidence strength is what turns Science knowledge into disciplined judgement. The learner observes before inferring, checks patterns before generalising, distinguishes association from cause and knows when a limitation changes the confidence of a conclusion.

The strongest answer is not the boldest claim. It is the strongest claim that the evidence can carry.

The Final Evidence Standard

The strongest Science response keeps every level distinct: observation is what was detected, pattern is what the data does, inference is what the evidence suggests, mechanism is why the pattern occurs, and conclusion is the bounded claim the whole investigation can support.

When checking, move upward through those levels deliberately. If the answer jumps from one observation to a universal conclusion, narrow it. If a well-controlled data set clearly supports a relationship, state it without unnecessary hesitation.

Evidence control is therefore calibrated confidence. The learner says no less than the data earns and no more than the design can carry.