Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

Bolt Performance Calibration — Two Assessments Can Correlate Strongly and Still Disagree About Individual Scores

Wait, What? Two Tests Can Correlate Almost Perfectly and Still Give Every Student a Different Score

Suppose Test A gives four students scores of 50, 60, 70 and 80. Test B gives the same students 60, 70, 80 and 90. The rank order is identical. The correlation is perfect. The scores do not agree: Test B is ten points higher for everyone.

This simple example exposes a common measurement mistake. Correlation asks whether two sets of scores move together. Agreement asks whether they are sufficiently close, on a sufficiently comparable scale, for the intended use.

Bolt therefore refuses the shortcut “these tests correlate highly, so we can use them interchangeably.” A high correlation can support one kind of comparability claim while leaving another completely unresolved.

Quick Answer

Owned Bolt calibration job: distinguish rank-order association from score agreement when comparing two assessments, so a strong correlation is not mistaken for evidence that individual scores are interchangeable.

Two assessments can correlate highly because students who do well on one tend to do well on the other. Yet one test may be systematically easier, use a different scale, measure a somewhat different construct or produce larger individual discrepancies at parts of the score range. If a school wants to substitute one score for another, much stronger comparability evidence is required.

Correlation Preserves Order, Not Equality

Correlation is useful. It can show that two assessments capture related performance. If high scorers on one test are generally high scorers on another, that is meaningful evidence.

But correlation is insensitive to some systematic differences. Add ten points to every score and the correlation does not change. Multiply every score by two and the correlation can still remain perfect. That means correlation alone cannot tell us that the score levels agree.

StudentTest ATest BDifference
A5060+10
B6070+10
C7080+10
D8090+10

The assessments rank students identically. If a scholarship threshold is 75, however, the two tests make different decisions for some students. That is why decision use determines how much comparability is needed.

Three Different Questions Hide Inside “Are These Tests Comparable?”

  • Association: Do students who score higher on Test A tend to score higher on Test B?
  • Agreement: Are the actual score values close enough for the intended individual interpretation?
  • Interchangeability: Can one test be substituted for the other without materially changing decisions or score meaning?

These are progressively stronger claims. A school should not leap from the first to the third.

Why Score Linking Exists

When different assessments are used for similar decisions, testing organisations sometimes conduct concordance or linking studies. These studies do more than report a correlation. They estimate relationships between score scales in a defined population, examine prediction error and state the conditions under which score comparisons are defensible.

ETS standards explicitly distinguish strict equating from broader linking. Alternate forms intended to be interchangeable require stronger equivalence than two different tests used for similar purposes. The score-user question is always: what kind of link has been established, for which population, and for what use?

Observable Evidence Patterns

  • High correlation, constant offset: rank order is preserved, but absolute score agreement is poor.
  • High correlation, widening differences: tests agree near the middle but diverge at high or low performance levels.
  • Moderate correlation, similar averages: group means may align while individual ranking differs substantially.
  • Strong group agreement, poor individual agreement: cohort comparisons may be defensible while individual substitution is not.
  • Good concordance in one population, weaker in another: linking relationships can be sample- and population-dependent.

What a High Correlation Can Support

A strong correlation can support the claim that two assessments order or differentiate students similarly in the population studied. That may be useful for research, broad ranking or evidence that the tests measure related constructs.

It can also be one component of a larger linking or validation study. But it is not sufficient evidence that a score of 80 on Test A means the same thing as 80 on Test B.

What a High Correlation Cannot Support by Itself

  • It cannot prove equal score scales.
  • It cannot prove equal difficulty.
  • It cannot prove that individual decisions will be the same.
  • It cannot prove that the tests measure exactly the same construct.
  • It cannot eliminate prediction error in a concordance table.
  • It cannot justify converting group averages using a table designed for individual scores unless the provider supports that use.
  • It cannot establish that the relationship will hold in a different population.

School–Teacher–Student Triad

School

If a school replaces one assessment with another, it should not use a high correlation as the entire migration plan. It should examine score scale, construct coverage, decision consistency, subgroup behaviour and any formal linking or concordance evidence.

Teacher or Coach

The teacher should ask whether disagreement is systematic or learner-specific. If Test B is always about ten points higher, the issue is different from a pattern in which some students rise twenty points and others fall twenty. The latter may indicate different construct emphasis or unstable individual agreement.

Student

The student should not interpret two different test scores as a contradiction in identity. Different assessments sample performance differently. The useful question is what each score was designed to represent and whether repeated evidence converges.

The Bolt Correlation-to-Agreement Calibration Protocol

  1. Name the intended use. Ranking, substitution, placement, growth tracking or research?
  2. Check construct similarity. Do the assessments actually target the same knowledge and skills?
  3. Check scale meaning. Are identical numbers intended to mean identical performance?
  4. Inspect correlation. How similarly do the scores order students?
  5. Inspect individual differences. Are there systematic offsets or large person-level discrepancies?
  6. Check formal linking. Equating, concordance or another linking method may be required.
  7. Check the population. A link established in one sample may not transport perfectly to another.
  8. Inspect decision consistency. Do thresholds or classifications change when the assessment changes?
  9. Use a common return task when needed. Fresh independent evidence can help interpret individual disagreement.

Worked Example: The School Replaces a Reading Test

A school adopts a new reading assessment. During transition, students take both old and new tests. The correlation between scores is .92. Leaders conclude that the tests are interchangeable and join the old and new scores into one longitudinal trend.

Bolt asks for the individual differences. The new test is systematically higher by eight scale points, and the discrepancy grows for students near the lower end because the new assessment has different text and vocabulary demands. Several intervention-threshold decisions change depending on which score is used.

The calibrated conclusion becomes: the assessments rank students similarly overall, but current evidence does not justify treating their raw scale scores as interchangeable for individual intervention decisions.

The school can still use both tests. It simply needs a more defensible linking or transition rule.

Group Comparability Can Be Better Than Individual Comparability

Averages can be relatively stable even when person-level differences are large. Two assessments might estimate a school mean similarly but disagree enough at the individual level to change placement decisions. The National Academy of Education’s work on large-scale assessment comparability treats individual and aggregate score comparability as distinct problems for exactly this reason.

Bolt therefore asks who the receiver of the score is. A result adequate for system monitoring may be inadequate for deciding what happens to one child.

Teacher–Student Dialogue

Student: “The two tests are highly correlated, so why did my scores look so different?”

Teacher: “Correlation tells us that students tend to keep a similar order. It does not tell us that the two numbers agree exactly for each person.”

Student: “Which one should I believe?”

Teacher: “We first check what each test measures and whether their scales were formally linked. Then we compare your performance on another common task before making a strong individual conclusion.”

For Parents

If a school says two tests are “basically the same because they correlate at .9,” ask whether the scores themselves agree closely enough for the decision being made. Ask whether there is a concordance or linking study, what its prediction error is and whether it was validated in a population similar to your child’s.

Also ask whether the issue concerns ranking, a broad school average or a decision about one student. Those uses require different levels of comparability.

How Do We Know?

The National Academy of Education volume Comparability of Large-Scale Educational Assessments: Issues and Recommendations treats comparability as a multi-dimensional validity problem and separates individual-score comparability from aggregate-score comparability.

The ETS Standards for Quality and Fairness state that when scores are meant to be comparable, appropriate equating or linking methods should be used, and the test-taker population and definition of comparability should be documented.

The 2025 ETS report Aligning Scores of Language Proficiency Tests: A Score Concordance Study Between IELTS Academic and TOEFL iBT illustrates modern score-concordance practice: similar-purpose tests require an empirical basis and explicit good-practice principles before individual score relationships are reported.

The joint ACT/SAT Concordance Guide explicitly warns users to consider prediction error, avoid decisions based solely on a concorded score and avoid converting aggregate scores with individual-score concordance tables.

Victoria Crisp’s review Exploring the Relationship between Validity and Comparability in Assessment argues that comparability concerns belong inside validation whenever score interpretations depend on comparison across assessment conditions.

Evidence Boundary

No single agreement statistic or correlation threshold guarantees interchangeability for every educational decision. Appropriate methods depend on score scale, construct, population and use. Bolt therefore does not invent a universal “correlation high enough” rule.

Assessment disagreement also does not diagnose a student’s motivation, anxiety, language disorder or any other clinical condition. It identifies an evidence discrepancy requiring better measurement.

Common Misconceptions

  • “Correlation of .9 means the scores are 90% the same.” Correlation is not a percentage of agreement.
  • “Perfect correlation means identical scores.” A constant or proportional difference can preserve perfect correlation.
  • “Same average means good individual agreement.” Person-level discrepancies can cancel in the mean.
  • “A concordance makes two tests identical.” Concordance supports a defined score relationship with error and boundaries.
  • “One test must be wrong if they disagree.” They may sample different aspects of performance or use different scales.

What Should Change Next?

If Bolt concludes, “The two assessments rank performance similarly, but this learner’s scores disagree beyond what the current link explains,” the next Bolt move is a common-condition performance or stronger linking evidence targeted to that discrepancy—not an automatic learner intervention.

If the fresh evidence identifies a genuine learning weakness, Bolt records that narrower finding but does not prescribe the learner operation. The measurement job remains the assessment discrepancy and what it justifies believing.

RFE: Did the next common evidence distinguish a scale/linking problem from a real learner-performance difference, and did school, teacher and student update only the claim the evidence can support?

Bolt Direction Graph

Assessment A + Assessment B → correlation check → individual agreement check → construct/scale/linking review → decision-consistency check → common-condition evidence where needed → calibrated comparability claim → school/teacher/student recalibration.

Useful neighbours: Bolt — A Raw Score and a Scaled Score Are Not the Same Kind of Number, Bolt — Before You Call It Improvement, Check Whether the Scores Are Comparable, and Bolt — A Teacher Can Rank Students Correctly and Still Misjudge Their Actual Level.