Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

Bolt Measurement Note 21 — A Test Question Can Behave Differently for Comparable Groups

Wait, What? A Question Can Be Perfectly Clear to One Group and Unusually Difficult for Another—Even When Overall Proficiency Is Similar

A test can have good overall reliability, sensible content and careful marking, yet still contain an item that behaves differently for different groups of learners. That does not automatically make the item biased. It does mean the item deserves investigation.

Educational measurement has a precise way of asking this question: after accounting for the proficiency the test is intended to measure, does a particular item still show different probabilities of success for different groups? This is the logic behind differential item functioning, usually shortened to DIF.

Quick Answer

Owned Bolt job: calibrate item-level fairness when a test question may function differently across comparable learner groups.

DIF is not a verdict that an item is unfair. It is a statistical flag that says: this item is behaving differently after we account for the measured proficiency, so inspect it. The difference may come from language, context, format, cultural familiarity, accessibility, unintended prerequisite knowledge, or another construct-irrelevant feature. It may also have an innocent explanation. The correct Bolt response is investigation, not accusation.

The Difference Between Group Performance and Item Fairness

Suppose Group A scores higher than Group B on a test. That group difference does not by itself prove the test is unfair. The groups may differ in preparation, opportunity to learn, language experience, prior instruction, or the target proficiency itself.

DIF asks a narrower question. Among learners who are comparable on the proficiency being measured, does one group still find a particular item systematically easier or harder?

That distinction matters because fairness cannot be inferred simply from equal average scores, and unfairness cannot be inferred simply from unequal average scores. A fair assessment can reveal real performance differences. An apparently neutral assessment can also contain local item behaviour that deserves scrutiny.

A Simple Example: The Mathematics Question That Quietly Became a Language Question

Imagine a mathematics item intended to measure proportional reasoning. The numerical relationship is straightforward, but the problem is embedded in an unfamiliar idiom or culturally specific situation. Two learners with similar proportional-reasoning proficiency may not have equal access to the wording.

If learners from one language-background group systematically underperform on that item relative to otherwise comparable learners, the item may show DIF. The next question is substantive: is the extra language demand part of the intended construct, or is it an unintended obstacle?

If language comprehension is intentionally part of the target, the difference may be legitimate. If the target is purely proportional reasoning, the same language feature may introduce construct-irrelevant difficulty.

DIF Is a Flag, Not a Guilty Verdict

This is one of the most important calibration rules in assessment fairness. Statistical DIF does not automatically establish bias.

  • An item can show DIF because the groups have genuinely different exposure to content that is legitimately part of the construct.
  • An item can show DIF because the matching variable does not perfectly represent the intended proficiency.
  • An item can show DIF because of sampling variation, especially with small groups.
  • An item can show DIF because of an unintended linguistic, cultural or accessibility feature.
  • An item can show little detectable DIF and still participate in a broader assessment design that has fairness problems elsewhere.

Therefore, good practice combines statistical detection with expert review, content analysis and a clear account of what the item is supposed to measure.

School, Teacher and Student: Three Different Fairness Responsibilities

School

Schools using common assessments should avoid assuming that one total reliability coefficient proves fairness for every subgroup and every item. When stakes rise, item review, accessibility review and subgroup performance analysis become more important. The aim is not to engineer identical outcomes. It is to reduce irrelevant barriers to the construct being assessed.

Teacher or Coach

A classroom teacher will rarely run formal DIF statistics. But the same reasoning can still improve judgement. If a cluster of otherwise capable students all miss one question, ask whether the item introduced an unintended demand before concluding that all of them share the same conceptual weakness.

Student

A student should not use fairness language to dismiss every difficult question. The disciplined question is narrower: was I unable to perform the intended skill, or was another demand unexpectedly carrying part of the task? That distinction can be tested using a cleaner version of the same target skill.

Competing Explanations When an Item Looks Uneven Across Groups

  • The item contains construct-irrelevant language or context.
  • The groups differ in legitimate opportunity to learn the tested content.
  • The groups differ in a prerequisite that is genuinely part of the construct.
  • The item format is more familiar to one group.
  • Accessibility or presentation features affect response opportunities.
  • The statistical flag is unstable because the subgroup sample is too small.
  • The overall matching score is not measuring exactly the same construct across groups.

None of these should be selected by intuition alone. The job is to gather enough evidence to distinguish them.

The Bolt Item-Fairness Calibration Protocol

  1. Name the intended construct. What exactly is this item supposed to measure?
  2. Identify the extra demands. Language, context, technology, reading load, visual interpretation, speed and background knowledge may all matter.
  3. Inspect subgroup patterns. Is the unusual item behaviour repeated or isolated?
  4. Condition on relevant proficiency where formal analysis is possible. Group averages alone are not DIF analysis.
  5. Review flagged items substantively. Ask subject experts, language/accessibility reviewers and assessment specialists what could explain the pattern.
  6. Check effect size, not only statistical significance. Very large samples can make tiny differences statistically detectable.
  7. Test revised or parallel items. If the unintended feature is removed, does the subgroup difference shrink while the intended construct remains measurable?
  8. Preserve uncertainty. A flag invites review; it does not prove motive, discrimination or invalidity.
  9. Recalibrate the assessment claim. Decide whether the item can remain, needs revision, should be excluded, or requires a narrower interpretation.

Worked Example: A Reading Item With a Familiarity Advantage

A reading-comprehension test uses a passage built around a particular leisure activity. The item is intended to measure inference, not knowledge of the activity. Learners with comparable overall reading performance show a persistent subgroup difference on one item.

The item is flagged. Reviewers then inspect the wording and realise that the correct inference depends heavily on knowing a convention associated with the activity. A revised item preserves the inferential structure but removes the background assumption.

If the subgroup difference reduces while the item continues to discriminate reading proficiency appropriately, the evidence supports the interpretation that the original item contained an unintended source of difficulty. That is how fairness review should work: not by guessing the learner, but by testing the measurement.

What This Does Not Mean

  • Different group outcomes do not automatically mean bias.
  • DIF does not automatically mean an item must be deleted.
  • No detected DIF does not prove the whole assessment is fair.
  • Fairness is not the same as making every question equally easy for everyone.
  • Classroom teachers do not need to become psychometricians. They do need the habit of asking whether an item is measuring what they think it is measuring.

How Do We Know?

The fifth edition of Educational Measurement includes a full chapter on Fairness in Educational Measurement, including differential item functioning, differential prediction and invariance across groups. A 2025 chapter from ETS, The Sociocultural Context of Educational Assessment, likewise treats validity, fairness, cultural differences and DIF as central to responsible educational assessment.

ETS has long used DIF as part of test-fairness review. Its report A Review of ETS Differential Item Functioning Assessment Procedures explains why statistical flagging rules, sample size and criterion choice matter. Recent research continues to refine interpretation. A 2025 study in Language Testing in Asia, Evaluating differential item functioning through effect size measures in logistic regression, argues for combining effect-size measures with statistical tests so practically important DIF is not confused with significance alone.

The evidence boundary matters: DIF identifies unusual conditional item behaviour. Establishing unfairness requires additional substantive evidence about the intended construct and plausible causes.

For Parents: Ask Whether the Question Measured the Intended Skill

If a child says, “I knew the topic but I did not understand what that question wanted,” do not immediately accept or dismiss the claim. Put the target skill into a cleaner task. If performance returns, the original item may have contained an additional demand worth inspecting. If performance still fails, the underlying skill remains the stronger explanation.

Fair assessment is not about protecting learners from difficulty. It is about making difficulty answerable to the thing we intended to measure.

Bolt Direction Graph

Item response → subgroup pattern → condition on intended proficiency → DIF flag → substantive review → inspect construct-irrelevant demand → test revision/parallel item → recalibrate item and score interpretation.

Useful neighbours: Bolt 03 — What Exactly Did This Test Measure?, Bolt Measurement Note 13 — A Test With Too Few Questions Can Misrepresent a Skill, and Bolt Measurement Note 18 — Crossing a Grade Boundary Does Not Create a Sudden Jump in Capability.