Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

Bolt Performance Calibration — A Mathematics or Science Test Can Become Partly a Language Test

Wait, What? A Student Can Know the Science and Still Lose the Question Before Reaching the Science

A learner sees a science item containing a long scenario, unfamiliar general academic vocabulary, a dense sentence structure and a final question asking for a simple scientific relationship. The learner struggles. The score is recorded under “science.”

Did the learner fail the science? Perhaps. But the performance may also reflect the language required to enter the task, understand its conditions and express the response. That does not make language irrelevant to science or mathematics. Scientific and mathematical practice genuinely use specialised language. The calibration problem is more precise: which language demands belong to the intended construct, and which ones add difficulty that the assessment did not mean to measure?

Quick Answer

Owned Bolt calibration job: determine whether linguistic demand in a content assessment is construct-relevant evidence of the intended mathematics or science performance, or construct-irrelevant difficulty that weakens the inference from score to subject knowledge.

A mathematics or science assessment cannot be language-free. Students must read symbols, labels, instructions, quantities, diagrams and often prose. Some language is essential to the subject. But educational research shows that unnecessary linguistic complexity can reduce some learners’ ability to demonstrate content knowledge, particularly learners still developing proficiency in the language of assessment.

Bolt therefore rejects both extremes: “language never matters in maths or science” and “any language difficulty is unfair contamination.” The validity question is whether the language belongs to the knowledge and practices the assessment intends to measure.

Language Is Part of the Performance Condition

Consider a mathematics word problem. A student may need to:

  • identify the quantities;
  • understand relational words such as “at least,” “remaining,” “per,” or “difference”;
  • separate relevant from irrelevant information;
  • translate the prose into a mathematical representation;
  • select a method;
  • calculate accurately;
  • communicate an answer in the requested form.

Some of that language processing may be inseparable from real mathematical problem solving. Other wording may simply make the question harder to parse without adding mathematical value. The assessment designer and score user need to know which is which.

Science Has the Same Problem — With an Important Twist

Science is communicated through explanations, evidence, models, specialised vocabulary, causal language and argument. A test of scientific explanation may legitimately require students to use language. Removing all linguistic demand could change the construct rather than improve it.

But a science item can also contain unnecessary syntactic complexity, confusing pronouns, long embedded clauses or excess reading that is not needed to demonstrate the target science. When those features disproportionately suppress performance, the score can begin to carry language variance that was not intended.

Observable Evidence Pattern

Task conditionObserved performancePossible interpretation
Dense worded science itemIncorrect / omittedScience weakness, language barrier, or both
Same concept with clearer wordingCorrectUnnecessary linguistic demand may have obscured knowledge
Diagram + short promptCorrect explanationConcept may be accessible through another representation
New scientific explanation taskStill weakLanguage alone does not explain the pattern

No one row settles the case. The changed conditions help discriminate among explanations.

Competing Interpretations of a Low Content Score

  • The learner genuinely lacks the mathematics or science knowledge.
  • The learner understands the content but cannot reliably access the question’s language.
  • The language demand is legitimately part of the intended construct.
  • The item contains unnecessary linguistic complexity that adds construct-irrelevant variance.
  • The learner knows the vocabulary but not the underlying concept.
  • The learner can explain orally but the written response format adds additional demand.
  • The difficulty comes from the representation, not language alone.
  • The result is an ordinary measurement fluctuation.

Bolt does not infer which explanation is true from the learner’s English-proficiency label, home language or one test score. It asks what changed evidence would separate the possibilities.

What Counts as Construct-Relevant Language?

There is no universal list. The answer depends on the educational claim.

  • If the target is numerical computation, a long narrative may be largely irrelevant.
  • If the target is mathematical modelling from a real-world situation, interpreting language may be essential.
  • If the target is recognising a scientific process, decorative prose may be irrelevant.
  • If the target is constructing a scientific explanation from evidence, disciplined use of language may be central to the construct.

This is why Avenia-Tapper and Llosa argue that linguistic complexity should not automatically be labelled construct-irrelevant simply because it is complex. The language should be judged against how knowledge is learned and used in the domain.

When Language Becomes Construct-Irrelevant Noise

Language becomes a validity threat when it makes performance depend on abilities that are not meant to be part of the score interpretation. Research on mathematics and science assessments has found associations between linguistic complexity and differential performance for English learners, especially when general academic vocabulary, amount of text or writing demand exceeds what the target content requires.

That does not mean simplifying every sentence always improves validity. Simplification can remove necessary disciplinary meaning or alter item difficulty in unintended ways. The right question is whether the modification preserves the construct while removing avoidable barriers.

School–Teacher–Student Triad

School

The school should examine whether content assessments use language appropriate to the intended construct and learner population. It should avoid treating a low subject score as pure evidence of subject weakness when the assessment requires substantial reading or writing beyond the target skill.

Teacher or Coach

The teacher should inspect the first point of failure. Did the learner misunderstand a scientific idea, misread a relational phrase, fail to decode the question command, or know the answer but struggle to express it? Those are different evidence patterns and require different next tests.

Student

The student should learn to describe the performance condition precisely: “I knew the formula but did not understand what ‘at most’ meant” is more useful than “I am bad at maths.” Likewise, “I understood the experiment but could not organise the written explanation” identifies a different calibration problem from not understanding the experiment.

The Bolt Language-Demand Calibration Protocol

  1. Name the construct. What mathematics or science knowledge or practice is the score intended to represent?
  2. Map the language job. What must the learner read, interpret, speak or write to access the task?
  3. Separate disciplinary language from unnecessary complexity. Do not simplify away the construct.
  4. Identify the first observable failure. Vocabulary, syntax, representation, concept, method, calculation, explanation or something else?
  5. Change one condition where possible. Use clearer wording, another representation or a matched item to test the language hypothesis.
  6. Keep content constant enough to compare. A new easier concept does not discriminate language from subject knowledge.
  7. Repeat across items. One language-sensitive item should not redefine the learner.
  8. Check delayed and transfer evidence. Does the content knowledge appear when the language context changes?
  9. Recalibrate the claim. State whether the evidence supports content weakness, language-access difficulty, both, or unresolved uncertainty.

Worked Example: The Fraction Problem That Was Really Two Problems

A student can add fractions accurately in symbolic exercises but misses a word problem asking for “the fraction of the remaining quantity.” The teacher initially concludes that transfer is weak.

The teacher asks the student to explain the phrase “remaining quantity.” The learner interprets it as the original total rather than what is left after the first change. A matched problem with a diagram is solved correctly.

The calibrated conclusion is not “the student cannot transfer fraction knowledge” and not “the problem was unfair.” It is: current symbolic fraction performance is stronger than performance on this linguistic representation; the phrase mapping needs confirmation across new word problems.

That narrower conclusion now tells Bolt what to measure next: a matched content performance under clearer language conditions, followed by a return performance that restores the ordinary disciplinary language demand. The learner operation, if any, is a separate ownership decision.

Worked Example: The Science Student Who Knows More Than the Written Score Shows

A Grade 5 learner gives strong oral explanations of evaporation during a practical lesson but repeatedly omits constructed-response science items. The school could conclude that the oral work was over-supported or that the child lacks written scientific understanding.

A better test compares response modes while preserving the concept. The learner chooses the correct evidence, identifies the causal relationship and can explain it orally, but struggles to construct the required written response in English. Research on Grade 5 science assessments has found that English writing demand can predict differential item functioning and that lower English proficiency can increase the odds of omitted constructed responses even after controlling for science proficiency.

The calibrated statement becomes narrower: the current written science score includes a substantial language-production demand; independent science understanding should be triangulated using additional evidence while written scientific communication is assessed as its own legitimate performance where required.

Teacher–Student Dialogue

Teacher: “You missed the question, but I need to know whether the science or the wording stopped you first.”

Student: “I did not know what ‘account for the observed change’ meant.”

Teacher: “Good. I’ll restate the instruction without changing the science. If you can now explain the evidence, that tells us something different from not understanding the concept.”

For Parents: Do Not Turn a Language–Content Discrepancy Into a Child Label

If a child knows mathematics orally or through diagrams but scores poorly on worded questions, ask where the performance breaks. The answer may involve language, the mathematical representation, the underlying concept, time pressure or several factors together.

For multilingual learners, avoid two opposite assumptions: “English is the whole problem” and “English should not matter at all.” The subject itself may require some academic language. The fair question is whether the assessment’s language is necessary for the intended claim.

Educational observations do not diagnose language disorder, dyslexia, ADHD, anxiety or another clinical condition. If a broader health or developmental concern exists, that belongs with appropriately qualified professionals.

How Do We Know?

The Institute of Education Sciences report Accommodations for English Language Learner Students: The Effect of Linguistic Modification of Math Test Item Sets explains the validity concern directly: students can be constrained in demonstrating mathematics knowledge when items measure factors beyond the intended mathematics construct. The study tested linguistic modification as one way to reduce unnecessary complexity without altering the mathematics target.

Wolf and Leon’s An Investigation of the Language Demands in Content Assessments for English Language Learners analysed 542 mathematics and science items from 11 assessments. General academic vocabulary and amount of language were among the features most strongly associated with differential item functioning for English learners, particularly those with lower English proficiency.

Noble and colleagues’ experimental study Targeted Linguistic Simplification of Science Test Items for English Learners tested original and modified Grade 5 science items with 310 English learners and 1,580 non-English learners. The study directly investigated whether specific linguistic features contributed construct-irrelevant variance to science scores.

The 2023 study English Learners and Constructed-Response Science Test Items: Challenges and Opportunities found that English writing demand was the strongest item-level predictor of differential functioning favouring non-English learners in the studied Grade 5 science assessment, while lower English proficiency was associated with much greater odds of omitting constructed responses even after controlling for science proficiency.

Importantly, Construct Relevant or Irrelevant? The Role of Linguistic Complexity in the Assessment of English Language Learners’ Science Knowledge provides a necessary boundary: complex language is not automatically irrelevant. Its relevance should be judged against the language actually used in the educational and disciplinary domain being assessed.

The broader validity and fairness framework is consistent with the Standards for Educational and Psychological Testing: interpretations should reflect the construct the assessment is intended to measure, and construct-irrelevant barriers should not silently drive consequential score differences.

Evidence Boundary

Much of the direct evidence concerns English learners in US assessment systems. The principle generalises cautiously: language demands can interact with content measurement, but the direction and size of the effect depend on age, language background, subject, item type, disciplinary practice and assessment purpose. Bolt does not assume that every multilingual learner is disadvantaged by every linguistically complex item, and it does not invent a universal “language-load correction.”

Common Misconceptions

  • “Maths is numbers, so language should not matter.” Many mathematical tasks require interpretation of language and representations.
  • “Any difficult wording is unfair.” Some disciplinary language is part of the construct.
  • “Simpler language always makes a test more valid.” Simplification can alter meaning or remove intended demand.
  • “A low science score for an English learner is really an English score.” That is too strong; content and language may both contribute.
  • “One clearer rewording proves the learner knows the whole topic.” Repeated and transfer evidence are still required.

What Should Change Next?

Suppose Bolt concludes: “The learner shows the target science relationship when linguistic access is clarified, but written disciplinary explanation remains weaker.” The next Bolt move is a matched return performance that preserves the science construct while varying only the language demand that remains uncertain.

If the return performance exposes a genuine language-production or content weakness, Bolt records that narrower finding. Any learner operation used to improve it is a separate ownership decision rather than part of a compulsory handoff.

RFE: Did the changed-condition and later disciplinary-language performances separate content knowledge from language demand well enough for school, teacher and student to update the claim without erasing legitimate scientific or mathematical language from the construct?

Bolt Direction Graph

Content construct → language demands → observed performance → construct-relevance check → changed-condition discrimination → calibrated content/language claim → disciplinary-language return performance → school/teacher/student recalibration.

Useful neighbours: Bolt — Multiple-Choice and Constructed-Response Scores Are Not Automatically Equivalent, Bolt — A Higher Score With an Accommodation Does Not Automatically Mean an Unfair Advantage, and Student/Studying Interface — Translation Tool Interface.