Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

Bolt Performance Calibration — Five Questions From One Passage Are Not Five Independent Pieces of Evidence

Wait, What? Five Correct Answers Can Sometimes Contain Less Than Five Independent Pieces of Evidence

A reading test presents one long passage followed by five questions. A student understands the passage unusually well and answers all five correctly.

The raw score records five correct responses. But the five answers are not completely independent events: they share the same passage, topic knowledge, interpretation and sometimes even information carried forward from earlier questions.

If a measurement model treats every item as though it were independent, it can overstate how much distinct information the cluster contributes about the learner.

Quick Answer

Owned Bolt calibration job: identify when clusters of questions built around a shared passage, graph, source, case or scenario create local dependence—so the apparent amount of independent performance evidence is larger than the true amount of distinct information.

Testlets are not bad test design. They are often necessary for reading comprehension, source analysis, science investigations and complex authentic tasks. The calibration problem appears when shared-stimulus dependence is ignored while interpreting score precision, subscores or learner profiles.

Why Shared Stimuli Create Dependence

Imagine five questions all depend on one source passage. Performance can become correlated for reasons beyond the target proficiency:

  • the learner already knows the passage topic well;
  • the learner misunderstands one central sentence that affects several answers;
  • an early question directs attention to information useful for later questions;
  • a graph or diagram is misread once, contaminating the whole cluster;
  • fatigue or interest changes across the shared scenario;
  • one response reveals or cues part of another response.

These shared influences mean the item responses can remain correlated even after accounting for the proficiency the test intends to measure.

Local Independence Is a Measurement Assumption

Many item-response models assume that once underlying proficiency is accounted for, the remaining item responses are conditionally independent. Testlets can violate that assumption because items share an extra source of variation.

The practical result can be information inflation: the test appears to know the learner more precisely than it actually does.

This does not mean the raw number of correct answers is false. It means the statistical interpretation of how much independent evidence those answers provide can be too optimistic.

Observable Signs of a Testlet Effect

  • A learner performs extremely well or poorly on all items tied to one passage but normally elsewhere.
  • Items within the same stimulus cluster correlate more strongly than expected from overall proficiency.
  • Prior topic familiarity appears to help or hurt an entire passage set.
  • Earlier items seem to cue later answers.
  • A subscore is dominated by a small number of shared stimuli rather than many independent content samples.
  • Model-based precision falls after testlet dependence is explicitly accounted for.

These patterns can reflect genuine skill too. The job is to determine how much of the evidence comes from the target construct and how much is shared across the cluster.

Competing Explanations for Five Correct Questions From One Passage

  • The learner genuinely has strong reading comprehension.
  • The learner had unusually strong prior knowledge of the passage topic.
  • The first correct interpretation made later questions easier.
  • The passage happened to match the learner’s vocabulary and background knowledge.
  • The testlet is well designed and the dependence is small.
  • The cluster is overcounting one narrow source of evidence.

Bolt does not downgrade the five correct answers automatically. It asks whether a fresh passage or scenario produces the same performance.

School–Teacher–Student Triad

School

When schools create common assessments, they should inspect how many independent stimuli support each important claim. Twenty questions can still provide narrow evidence if fifteen depend on only three passages or scenarios.

Teacher or Coach

If a learner appears weak on one passage cluster, the teacher should avoid diagnosing five separate weaknesses from five linked errors. A new stimulus testing the same reasoning can reveal whether the problem travels.

Student

The student should understand that one excellent passage set is encouraging but not the same as broad secure performance. Likewise, one disastrous cluster should not become a global identity judgement if the difficulty was tied to that source.

The Bolt Shared-Stimulus Calibration Protocol

  1. Map the stimulus structure. Which items share passages, graphs, cases or scenarios?
  2. Count independent evidence sources, not only item count.
  3. Inspect whether one misunderstanding can propagate.
  4. Check prior-knowledge sensitivity where relevant.
  5. Inspect item-order effects inside the cluster. Earlier responses may affect later ones.
  6. Use psychometric models that account for testlet effects when stakes justify it.
  7. Do not overinterpret tiny testlet-based subscores.
  8. Use a fresh independent stimulus. Preserve the target skill while changing the shared context.
  9. Compare whether the performance pattern returns.
  10. Recalibrate the breadth and precision of the claim.

Worked Example: Five Questions, One Misread Graph

A Science assessment presents one experimental graph followed by five questions. A student misreads the x-axis scale. That single interpretation error leads to four wrong answers.

A raw item analysis could conclude that the student is weak in trend description, comparison, interpolation and causal interpretation.

Bolt asks whether those are four independent weaknesses. The teacher gives a fresh graph with the same reasoning demands and a different axis structure. The learner answers four of five correctly.

The calibrated conclusion becomes narrower: the original cluster amplified one graph-reading error across several linked items; broad scientific reasoning is stronger than the five-item error pattern suggested.

How Do We Know?

A 2025 Journal of Educational Measurement paper, Parametric Bootstrap Mantel–Haenszel Statistic for Aggregated Testlet Effects, describes how shared stems induce correlations among item responses and why ignoring local dependence can bias item, ability and uncertainty estimates.

The 2025 Psychometrika paper Adjusting for Information Inflation Due to Local Dependency in Moderately Large Item Clusters focuses directly on the key Bolt issue: treating dependent clustered items as independent can overestimate the information contained in the test.

Modeling Directional Testlet Effects on Multiple Open-Ended Questions, published in 2025, extends the problem to open-ended items and shows that earlier responses within a testlet can positively or negatively influence later responses. Ignoring such directional dependence can distort reliability estimates.

ETS has studied testlet dependence for decades. Its report Using a New Statistical Model for Testlets to Score TOEFL found that assuming conditional independence in passage-based clusters could substantially overestimate test information.

Evidence boundary: local dependence is not automatically large enough to matter in every testlet. Shared stimuli are often essential for authentic assessment. The size and consequence of a testlet effect depend on item design, stimulus, population, scoring model and intended use.

Common Misconceptions

  • “Questions from one passage are redundant.” No. They can measure several legitimate aspects of comprehension.
  • “Five correct answers only count as one.” No. The issue is dependence, not deleting legitimate item evidence.
  • “Local dependence means the test is invalid.” Testlet models can account for it, and many strong assessments use shared stimuli.
  • “This is just content sampling.” Content sampling asks how broadly the domain is represented; local dependence asks whether supposedly separate item evidence shares extra variance.
  • “A poor passage cluster proves a broad weakness.” A fresh independent stimulus is needed before enlarging the claim.

What Should Change Next?

When one shared-stimulus cluster drives the diagnosis, Bolt asks for another independent stimulus that preserves the target reasoning but removes the original passage, graph or case. The aim is to see whether the weakness travels.

RFE: Did the learner reproduce the same performance pattern on an independent stimulus, or did the new evidence show that the original cluster had overcounted one narrow source of difficulty or strength?

Bolt Direction Graph

Shared stimulus → clustered item responses → local-dependence check → information/precision adjustment → independent stimulus return → calibrated breadth-of-evidence claim.

Useful neighbours: A Test With Too Few Questions Can Misrepresent a Skill, Question Order Can Change the Performance You Measure, and A Total Score Can Hide a Changing Skill Profile.