Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

PSLE Science Reality Lab Vol No.081 | “12 Failures Versus 8” — Were the Same Number of Items Tested?

PSLE-SCI-REALITY-0081

Wait, What? Twelve failures can be better evidence than eight failures—if far more items were tested.

A product-comparison infographic prints two large numbers:

  • Product A: 12 failures
  • Product B: 8 failures

The obvious reaction is that Product A performed worse. Twelve is larger than eight.

Then you find the test sizes. Product A was tested on 1,200 items. Product B was tested on 100 items. Under the same simplified test conditions, A failed 12 out of 1,200, while B failed 8 out of 100.

Now the scientific comparison changes. A’s failure fraction is 12/1,200 = 1%. B’s is 8/100 = 8%. The larger raw count came from a much larger tested group.

Reality Lab Vol No.081 teaches one real-world evidence-transfer job: when a scientific or product claim compares event counts from groups of different sizes, find the denominator before deciding which result represents the larger proportion.

Quick Answer

  1. Read the count. How many failures, successes, sightings or events occurred?
  2. Find the denominator. Out of how many tested items or opportunities?
  3. Check the counting unit. Are we counting failed items, failure events or repeated tests?
  4. Convert to a comparable fraction or percentage when appropriate. Compare 12/1,200 with 8/100, not 12 with 8 alone.
  5. Check exposure and method. Comparable denominators do not help if the groups were tested differently.
  6. Limit the conclusion. A lower observed failure proportion in this sample is not automatic proof of universal superiority.

Reality Lab habit: A count tells you how many. A denominator tells you how often out of the opportunities that existed.

The Owned Learner Job — and the Boundary

This page does not become the general owner of fractions, percentages, sample size or unequal-group comparison. Those skills already exist elsewhere. Vol No.081 applies them to one durable evidence object: a real-world reliability or test graphic that compares raw failure counts while hiding or downplaying unequal numbers tested.

Vol No.005 starts with a percentage claim and asks who sits underneath the denominator. Vol No.081 starts with raw counts and asks whether comparing the counts without denominators reverses the scientific impression.

Original Reality Lab Case: The Hinge-Cycle Poster

This is an original composite teaching case. No real product, manufacturer or advertisement is being criticised.

Two simple hinge designs are subjected to the same classroom reliability test. Each tested item is opened and closed for the same number of cycles. A failure means the hinge can no longer complete the test motion.

DesignItems testedItems that failedObserved failure fraction
A1,2001212/1,200 = 1%
B10088/100 = 8%

A poster that prints only “12 failures” and “8 failures” invites the wrong comparison. Product A has more failure events in total because far more A items were tested. Relative to the number tested, A had the lower observed failure fraction.

Observed, Claimed and Inferred

LayerWhat can be said
Observed12 of 1,200 A items failed; 8 of 100 B items failed.
ClaimedA is less reliable because 12 is greater than 8.
InferredThe raw number of failures is being compared as though the same number of items had been tested in both groups.

The count itself is not wrong. The comparison is incomplete.

The Denominator Is Part of the Scientific Quantity

“12 failures” is a count. “12 failures out of 1,200 tested items” is a proportion-bearing result. The second statement carries information about opportunity.

If two groups are the same size and receive the same test, raw failure counts can be compared directly. If group sizes differ, the larger group has more chances to accumulate failures even when its underlying failure tendency is lower.

Representation Check: Big Numbers Attract the Eye

Infographics often display counts in large type because whole numbers are easy to read: “12 failures”, “8 complaints”, “430 detections”, “27 breakdowns”. The denominator may appear in tiny text—or disappear entirely.

A strong scientific representation keeps numerator and denominator together. For example:

  • 12 failures among 1,200 tested items (1%)
  • 8 failures among 100 tested items (8%)

Now the reader can see both the amount of evidence and the relative frequency.

Comparison Check: Use Like With Like

Before converting counts to percentages, make sure the denominators mean comparable things.

  • Same type of item?
  • Same definition of failure?
  • Same test duration or number of cycles?
  • Same environmental conditions?
  • Same inspection method?
  • Same rule for removing invalid tests?

If Product A is tested for 10 cycles and Product B for 10,000 cycles, dividing by the number of items alone would still miss an important difference in exposure. The denominator must fit the event being counted.

Counting Unit Check: Failed Items or Failure Events?

Suppose one machine can fail several times and be repaired between runs. “12 failures” could mean 12 different machines each failed once, or one machine failed repeatedly. Those are not the same evidence structure.

Ask what one count represents. The denominator might be items, trials, hours, cycles, visits or observation opportunities. Keep that unit stable before comparing.

Method and Variable Check

Equalising the denominator solves only one problem. You still need a fair scientific comparison.

  • Were the same loads applied?
  • Were testing temperatures comparable?
  • Were the same failure criteria used?
  • Were testers blinded to the design where subjective judgement was possible?
  • Did every item receive the same planned exposure?
  • Were missing or interrupted tests reported?

A mathematically correct percentage from an unfair test does not become strong evidence merely because it includes a denominator.

Source and Provenance Check

A reliability graphic should let you trace the result backward: claimed rate → failures counted → number tested → test method → individual observations. If the numerator is easy to find but the denominator is hidden, the claim is harder to evaluate.

Also ask whether all tested items were included. A denominator can shrink if inconvenient tests are silently removed.

Worked Case 1: More Failures, Lower Failure Fraction

Group X has 20 failures among 2,000 items. Group Y has 10 failures among 100 items.

X: 20/2,000 = 1%. Y: 10/100 = 10%. Raw counts make X look worse, but the observed fraction is much lower for X.

Worked Case 2: Same Count, Different Meaning

Two groups each record 5 failures. Group M tested 50 items; Group N tested 500.

M: 10%. N: 1%. Equal counts do not mean equal rates when opportunities differ.

Worked Case 3: Equal Denominators

Product C and Product D each have 200 tested items under the same method. C has 6 failures and D has 10.

Here the equal denominator means the raw counts already preserve the same comparison: 3% versus 5%. You can still calculate the fractions, but the ranking does not change.

Worked Case 4: Different Test Exposure

Two lamp designs each have 100 units. Design A is run for 100 hours per unit and records 4 failures. Design B is run for 1,000 hours per unit and records 6 failures.

Comparing 4/100 with 6/100 alone ignores the much longer exposure for B. The correct reliability measure may need time or operating cycles. The simple Primary Science lesson is to notice that the “opportunity to fail” is not matched.

Worked Case 5: Missing Denominator

A social post says, “Brand X had only three failures in testing.” Without the number tested, what can you conclude?

Very little about the failure proportion. Three failures out of 30 is 10%; three out of 30,000 is 0.01%. The numerator alone leaves a huge range of possible interpretations.

Alternative Explanations to Keep Alive

  • The groups may differ in size.
  • The groups may have different exposure times.
  • The failure definition may differ.
  • One group may include harder test conditions.
  • Some missing tests may not be included in the denominator.
  • One large sample may reveal rare failures that a small sample simply did not have enough opportunities to show.

These possibilities do not tell you the final answer. They tell you which details are needed before the count becomes a fair comparison.

Sample Size Matters in Two Different Ways

First, sample size sets the denominator. Second, a larger sample usually gives more opportunities to observe uncommon events. That means a large study can record more failures in total while giving a more stable estimate of a low failure proportion.

Do not turn this into the opposite mistake: “the larger sample must be correct.” A large biased or unfair sample can still mislead. Size improves some evidence problems; it does not repair all of them.

What Evidence Would Strengthen the Reliability Comparison?

  • Numerator and denominator for each group.
  • A clear definition of one failure.
  • Comparable test duration or exposure.
  • Same test method and environmental conditions.
  • Transparent handling of missing or invalid trials.
  • Individual or batch information showing the sample was not selectively chosen.
  • Replication or additional testing under relevant conditions.

What Would Weaken It?

  • Only raw failure counts are displayed.
  • Group sizes are omitted.
  • Different exposure lengths are hidden.
  • Different definitions of “failure” are used.
  • The smaller count comes from a much smaller sample.
  • Failed or incomplete tests are excluded without explanation.
  • The sample is used to make a universal product claim without broader evidence.

How Far Can the Conclusion Travel?

If A shows a 1% observed failure fraction and B an 8% fraction in one well-matched test, you have evidence about those samples and conditions. You do not yet know the exact long-term failure rate in every environment, every production batch or every future use.

Generalisation requires evidence that the tested items and conditions represent the intended claim.

Model and Measurement Limits

Real reliability analysis can use confidence intervals, survival methods, censoring and exposure-adjusted rates. Primary learners do not need those tools here. The aim is more fundamental: do not compare event counts as though the opportunity to produce those events were automatically equal.

A percentage is also not magic. If conditions differ or the denominator is the wrong unit, the percentage may still answer the wrong question.

PSLE-Style Transfer Case

Two groups of seedlings are observed for the same period under the same conditions. In Group P, 9 of 300 seedlings wilt. In Group Q, 5 of 50 wilt. A student says Group P was affected more because nine seedlings wilted compared with five.

Evaluation: The groups have different sizes. Group P has 9/300 = 3% wilted, while Group Q has 5/50 = 10%. The raw count is larger in P, but the proportion affected is larger in Q. The learner should compare the fraction of each group, while still checking whether other test conditions were kept comparable.

Tempting Reasoning That Fails

  • “12 is bigger than 8, so A is worse.” Not until the number tested is known.
  • “The smaller percentage proves B is always better.” It supports a sample-level comparison under the tested conditions, not a universal claim.
  • “A bigger sample is unfair because it has more chances to fail.” A bigger sample can be valuable; compare proportions or suitable rates rather than raw totals.
  • “Once I calculate a percentage, the comparison is fair.” Test conditions and exposure must still match.
  • “Zero failures means the true failure rate is zero.” A finite sample can simply fail to observe a rare event.

Explained Practice

Practice A: A records 6 failures among 600 items; B records 4 among 80. Which observed failure proportion is lower? A: 1% versus 5%.

Practice B: A and B each test 100 items. A has 7 failures and B has 3. Can raw counts be compared directly for the sample-level ranking? Yes, because the denominators match, assuming the test conditions do too.

Practice C: A has 2 failures in 50 items tested for one day. B has 3 failures in 50 items tested for one month. Is 2/50 versus 3/50 enough for a fair long-term reliability comparison? No. Exposure differs greatly.

Delayed Independent Return: The O-U-T-O-F Check

  1. O — Outcome: What event is being counted?
  2. U — Unit: Is one count an item, event, trial or time period?
  3. T — Tested total: What is the denominator?
  4. O — Opportunity: Did both groups have comparable chances to produce the event?
  5. F — Fraction: What proportion or suitable rate should be compared?

Try this later on a new infographic. If you automatically ask “out of how many?” before judging the larger count, the scientific habit has transferred.

Parent and Tutor Teaching Guide

Use two bowls of counters. Bowl A has 100 counters with 8 red “failures”. Bowl B has 20 counters with 4 red failures. Ask which bowl has more red counters. The answer is A. Then ask which bowl has the larger fraction red. The answer is B: 20% versus 8%.

Next make both bowls the same size. The learner should notice that raw counts become directly comparable when opportunities match. Then change the exposure: tell them one bowl represents items tested for ten times longer. This reveals the next layer—sometimes the simple denominator is not enough.

The teaching goal is not to turn Primary Science into statistics. It is to make the learner suspicious of a count that arrives without its scientific opportunity structure.

Authoritative Sources

NIST’s quality-control guidance distinguishes counts of nonconforming items from proportions and explicitly treats the number of defectives relative to the number observed. The statistical methods are beyond Primary Science, but the evidence principle is not: the denominator is essential when group sizes differ. SEAB’s current PSLE Science objectives include interpreting and analysing information and evaluating observations, information and methods, which is exactly the reasoning this transfer case practises.

The Quiet Return

Large numbers attract attention. Small numbers feel safe.

Science asks a quieter question before deciding what either number means:

Out of how many opportunities?

Once you ask that, a count becomes evidence instead of decoration.