Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

Bolt Performance Calibration — Negative Marking Can Measure Risk Strategy as Well as Knowledge

Wait, What? Two Students With the Same Knowledge Can Earn Different Scores Because One Is More Willing to Risk a Penalty

A multiple-choice test gives one mark for a correct answer and deducts part of a mark for a wrong answer. Two students reach the same uncertain item. Both can eliminate two options but are not sure which of the remaining two is correct. One answers. The other leaves it blank.

If the first student is correct, the score rises. If wrong, the score falls. The second student receives neither gain nor penalty. The scoring rule is no longer measuring only whether students know the answer; it also changes the decision they must make under uncertainty.

Bolt therefore asks: how much of the observed score reflects knowledge, partial knowledge, guessing behaviour and risk strategy under the declared scoring rule?

Quick Answer

Owned Bolt calibration job: interpret negatively marked or formula-scored multiple-choice performance when penalties for wrong answers alter guessing and omission behaviour, so the final score is not automatically treated as a pure measure of subject knowledge.

Negative marking is often introduced to discourage blind guessing. Under some assumptions it can improve aspects of score reliability or make random guessing less profitable. But decades of educational-measurement research show that scoring rules also influence test-taking strategy. Students differ in willingness to answer uncertain items, and partial knowledge is not the same as random guessing.

The correct calibration is therefore not “negative marking is unfair” or “negative marking is more accurate.” The scoring rule has benefits and costs that must be evaluated against the construct and decision.

Number-Right and Negative Marking Create Different Decision Environments

Under conventional number-right scoring, a correct answer earns credit while an incorrect or omitted answer commonly earns zero. If there is no penalty for a wrong answer, students often have an incentive to respond even when uncertain.

Under formula or negative scoring, an incorrect answer carries a deduction intended to offset some advantage from random guessing. That changes the student’s expected-value calculation. An item is no longer simply “Do I know this?” It becomes “How certain am I, how many options can I eliminate, and is answering worth the risk under this scoring rule?”

Partial Knowledge Breaks the Simple Guessing Model

Classic correction-for-guessing formulas often begin from a simplified idea: when a student does not know an answer, they guess randomly. Real students frequently know something. They may eliminate one distractor, recognise that two options are impossible, remember part of a rule or hold a misconception that makes one distractor especially attractive.

Frederic Lord’s foundational work on formula scoring argued explicitly that the assumption “know or random guess” is usually implausible. That matters because the scoring rule interacts with partial knowledge. A student who can narrow four options to two has different evidence from a student choosing blindly among four, even if both are labelled “uncertain.”

Observable Evidence Pattern

Student stateLikely response under no penaltyPossible response under negative markingWhat Bolt notices
Knows answerAnswerAnswerStable evidence
Can eliminate two of four optionsUsually answerAnswer or omit depending on strategyPartial knowledge + risk decision
Pure guessOften answerMore likely omitPenalty may reduce blind guessing
Highly risk-averseAnswer uncertain itemsMay omit even with useful partial knowledgeStrategy can suppress score

The same content knowledge can therefore produce different observed scores when scoring rules change.

What Negative Marking Can Support

When well designed and clearly communicated, negative marking can reduce the incentive for completely random answering. It can make omission behaviour informative about uncertainty and, in some settings, has been associated with reliability benefits.

If the intended construct includes calibrated decision making under uncertainty, that behavioural component may even be relevant. Some professional contexts genuinely require a person to know when evidence is insufficient to act.

What Negative Marking Cannot Support by Itself

  • It cannot guarantee that every omitted item represents zero knowledge.
  • It cannot guarantee that every attempted uncertain item is blind guessing.
  • It cannot remove all guessing-related variance.
  • It cannot prove that two students with equal scores used equal knowledge and equal risk strategies.
  • It cannot justify comparing a negatively marked score directly with a number-right score as though only the arithmetic changed.
  • It cannot diagnose a stable personality trait from one test-taking pattern.
  • It cannot tell us whether the scoring rule improves validity for the specific educational decision without evidence.

Risk Aversion Can Become Construct-Irrelevant Variance

If the test is supposed to measure mathematics knowledge, but two equally knowledgeable students receive different scores mainly because one is more cautious about penalties, risk preference is entering the score without being an intended mathematics construct.

Research comparing negative marking with alternatives has repeatedly raised this concern. A study of “don’t know” responses in medical education found that formula scoring could disadvantage students who were less willing to guess, while number-right scoring had lower reliability in that setting. The result is not a winner-takes-all verdict; it exposes a trade-off between psychometric and behavioural properties.

A later Rasch-model comparison likewise concluded that scoring-method choice should consider not only psychometric characteristics but also self-directed test-taking strategies and metacognitive behaviour.

The Scoring Rule Can Change Behaviour Before It Changes the Score

The most important causal order is often:

scoring rule → response strategy → omissions/attempts → observed score.

If the scoring rule changes, schools should not interpret a score difference as a knowledge difference until they have checked how the response pattern changed. A student may answer fewer items under penalties while accuracy on attempted items rises. Another may keep answering almost everything and accept more deductions. The total can conceal those pathways.

School–Teacher–Student Triad

School

The school should decide why negative marking is being used and whether the penalty structure supports the intended interpretation. If the rule changes from year to year, score comparisons should be treated as a changed measurement condition.

Teacher or Coach

The teacher should inspect correct, incorrect and omitted responses separately. Two identical totals can come from very different profiles: high certainty with few attempts, broad attempts with many penalties, or strong partial knowledge combined with cautious omission.

Student

The student should understand the scoring rule before the performance and should not turn caution or willingness to attempt into identity. Prediction can help: before answering an uncertain item, what probability does the learner assign to being correct, and does the later outcome reveal systematic over- or under-caution?

The Bolt Negative-Marking Calibration Protocol

  1. Name the scoring rule. What is earned for correct, wrong and omitted responses?
  2. Name the intended construct. Knowledge only, or knowledge plus calibrated decision making under uncertainty?
  3. Predict behavioural effects. Will the rule change omissions, guessing or time use?
  4. Inspect response categories separately. Correct, wrong and blank carry different evidence.
  5. Check partial knowledge. Avoid treating every uncertain response as a random guess.
  6. Compare like with like. Do not compare totals across different penalty rules without calibration.
  7. Use confidence evidence cautiously. Prediction-before-answer can help distinguish knowledge from strategy, but confidence is not itself correctness.
  8. Collect an independent receipt. A short non-penalised or constructed-response task can test whether the apparent weakness survives a changed scoring condition.
  9. Recalibrate the claim. State whether the score reflects weak knowledge, cautious strategy, guessing, or unresolved mixture.

Worked Example: Same Knowledge, Different Penalty Behaviour

Two students each know 30 items securely and can narrow another 10 four-option items to two plausible answers. On a 50-item test, wrong answers lose one-third of a mark.

Student A attempts most of the uncertain items. Student B leaves nearly all of them blank. Their final scores differ. The teacher could conclude that A knows more. Bolt asks for another performance in which the same target knowledge is sampled with a different response format or without a guessing penalty.

If the students then perform similarly, the original difference was at least partly strategy-dependent. If A remains stronger across formats, the evidence supports a broader capability difference. The changed condition discriminates the explanation.

Teacher–Student Dialogue

Teacher: “You left twelve questions blank. That does not automatically mean you knew nothing about them.”

Student: “I could usually remove one or two answers, but I did not want to lose marks.”

Teacher: “That matters. Your score reflects both what you knew and how you handled uncertainty under this penalty rule. We need another piece of evidence before I call those topics weak.”

For Parents

If your child’s score drops when negative marking is introduced, ask whether the number of omitted questions changed and whether accuracy on attempted questions changed. A lower total can reflect weaker knowledge, greater caution or both.

Do not tell a cautious child to “just guess everything” unless that advice is mathematically correct under the actual scoring rule and appropriate to the assessment purpose. Likewise, do not interpret many attempts as bravery or many omissions as lack of confidence. The evidence needs a declared scoring context.

How Do We Know?

Frederic Lord’s classic Formula Scoring and Number-Right Scoring challenged the simplistic assumption that students either know an answer or guess randomly and examined the implications for scoring multiple-choice tests.

Rowley and Traub’s Formula Scoring, Number-Right Scoring, and Test-Taking Strategy directly links scoring method with test-taking behaviour, supporting the principle that the rule can alter the performance being measured.

The review Scoring Methods for Multiple Choice Assessment in Higher Education — Is It Still a Matter of Number Right Scoring or Negative Marking? surveys strengths, weaknesses and alternatives across the scoring literature rather than treating either method as universally superior.

A medical-education study, The Effect of a “Don’t Know” Option on Test Scores: Number-Right and Formula Scoring Compared, found a bias under formula scoring associated with willingness to guess, while also finding higher reliability under formula scoring in that setting. This is a useful example of competing measurement properties.

A later comparison, Comparison of Formula and Number-Right Scoring in Undergraduate Medical Training: A Rasch Model Analysis, found stronger psychometric properties for number-right scoring in the studied tests and argued that scoring decisions should consider test-taking strategies as well as psychometrics.

The PLOS ONE study Elimination Testing With Adapted Scoring Reduces Guessing and Anxiety in Multiple-Choice Assessments compared an alternative scoring approach with negative marking and explicitly treated risk aversion as a potentially irrelevant student characteristic in score production.

Evidence Boundary

Much direct scoring-rule research comes from higher education and medical education. The behavioural mechanism can inform school assessment, but age, stakes, item format and student understanding of the rule matter. There is no universal negative-marking penalty that removes guessing without introducing other trade-offs.

Risk behaviour observed on one test is not a clinical diagnosis or a stable personality verdict. Bolt keeps the interpretation tied to the scoring condition and repeated educational evidence.

Common Misconceptions

  • “Negative marking removes guessing.” It changes the incentive; guessing and partial knowledge can remain.
  • “A blank means the student knew nothing.” The student may have partial knowledge but judge the penalty too risky.
  • “A wrong attempted answer proves overconfidence.” One item is insufficient evidence.
  • “Formula scoring is always more accurate.” Research shows trade-offs and context dependence.
  • “Same subject + same questions means comparable scores after a scoring-rule change.” The response strategy itself may change.

What Should Change Next?

Suppose Bolt concludes: “The learner’s low score is partly driven by high omission under negative marking; topic knowledge remains underdetermined.” The next Bolt move is a short changed-condition performance that samples the same content under a response format or scoring rule that reduces the penalty-strategy confound.

If the changed-condition result confirms a genuine knowledge weakness, Bolt records that narrower claim. Any learner operation used to improve the weakness is a separate ownership decision rather than part of a compulsory route from this score.

RFE: Did the changed-condition performance separate subject knowledge from penalty-driven response strategy well enough for school, teacher and student to update the interpretation without overstating either weakness?

Bolt Direction Graph

Knowledge + uncertainty → scoring rule → answer/omit strategy → observed score → response-pattern/risk check → changed-condition performance → calibrated knowledge-versus-strategy claim → school/teacher/student recalibration.

Useful neighbours: Bolt — A Correct Multiple-Choice Answer Does Not Always Mean the Student Knew It, Bolt — A Blank Response Is Missing Evidence, Not Automatically Zero Capability, and Bolt — Multiple-Choice and Constructed-Response Scores Are Not Automatically Equivalent.