Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

How to Learn Advanced English Vocabulary (Chinese Edition) | Lesson No.044 | Master Accuracy, Precision, Bias, Error, Reliability, Repeatability and Reproducibility Without Treating Every Consistent Number as Correct | 第044课:掌握准确度、精密度、偏差、误差、可靠性、重复性与再现性,避免把“稳定一致”误当“正确”

Series ID: EDKS-ADV-VOC-ZH-0044 · Advanced English Vocabulary (Chinese Edition) · Lesson No.044 · C1 → C2 · 简体中文辅助

Consistency tells you whether measurements agree with each other. Accuracy asks whether they agree with the accepted reference.
一致性告诉你测量彼此是否相符;准确度问它们是否接近参考值。

A measuring system can return 9.80, 9.80, 9.81 and 9.79 every time when the accepted reference is 10.00. Those readings are tightly grouped—high precision/repeatability—but systematically low. This is the core distinction: repeated agreement is not proof of correctness.

NIST’s current accuracy glossary defines accuracy as closeness of a test result to an accepted reference value and notes that a set of results combines random components with systematic error or bias. NIST also distinguishes repeatability—agreement under the same conditions—from reproducibility—agreement under changed conditions. This lesson builds a public language system around those differences.

A precise instrument can be consistently wrong.

一个很精密的仪器,也可以稳定地测错。

This lesson owns measurement-quality vocabulary: accuracy, precision, trueness, bias, error, random/systematic error, reliability, repeatability, reproducibility, calibration, uncertainty, stability, agreement and validity. It coordinates with Lesson 039 on spread, Lesson 041 on correlation/agreement, and Lesson 043 on forecast accuracy.

Part I — Build the measurement-quality map | 第一部分:建立“测量质量地图”

1. Measurand is the quantity intended to be measured | measurand

Metrology uses measurand for the quantity intended to be measured. Accuracy and error make no sense until you know what quantity the measurement is meant to represent.

2. Reference value is the comparison anchor | reference value

Accuracy requires an accepted/reference value. In practice the true value may be unknown, so a calibrated standard or accepted reference stands in for it.

3. Accuracy asks closeness to reference | accuracy

NIST defines accuracy as closeness of agreement between a test result and accepted reference. For a set of results, accuracy reflects both random and systematic components.

4. Accurate does not mean repeatable automatically | accurate

Individual readings can scatter around the reference: the average may be good while precision is poor. Accuracy and consistency answer different questions.

5. Precision asks closeness among repeated results | precision

NIST/ISO terminology treats precision as closeness of agreement among independent results obtained under specified conditions. It says nothing by itself about closeness to the reference value.

6. Precise can still be inaccurate | precise but wrong

If every reading is tightly clustered at 9.80 when reference is 10.00, the system is precise but biased/inaccurate.

7. Imprecise measurements can average near the truth | imprecise

Readings can scatter widely above and below the reference while their mean is close. The process may have low bias but poor precision.

8. Trueness concerns systematic closeness of mean to reference | trueness

Modern metrology distinguishes trueness from precision. Trueness concerns closeness of the average of repeated measurements to a reference; precision concerns mutual agreement among measurements.

9. Bias quantifies systematic offset | bias

NIST describes measurement bias as the difference between an average measurement result and the true/target/reference value. A positive/negative bias indicates systematic direction.

10. Bias is not any mistake | systematic bias

One random bad reading is not “bias.” Bias is persistent/systematic tendency across measurements or estimates under a defined process.

11. Error is difference between measured and reference value | error

Measurement error is the difference between a result and a reference/true value under a sign convention. The exact true error is often unknowable because the true value itself is imperfectly known.

12. Random error creates scatter | random error

Random error changes unpredictably across repeated measurements and contributes to imprecision. Averaging independent repeated observations can reduce its effect on the mean.

13. Systematic error creates directional offset | systematic error

Systematic error shifts measurements in a consistent direction. Repeating the same biased measurement does not make the bias disappear.

14. Averaging reduces random error, not systematic bias | averaging

Collecting many readings can improve precision of the mean, but if the instrument is systematically 0.2 units low, the average remains about 0.2 low unless bias is corrected.

15. Repeatability is same-condition agreement | repeatability

NIST defines repeatability around repeated results under the same person, instrument, procedure, place and similar time conditions. It is a narrow precision condition.

16. Reproducibility tests changed-condition agreement | reproducibility

NIST defines reproducibility as agreement under changed conditions such as operator, instrument, laboratory, location or time. The changed conditions should be specified.

17. Repeatable is easier than reproducible | same vs changed conditions

A method can work consistently for one operator on one device yet vary across laboratories. Repeatability does not guarantee reproducibility.

18. Reproducible does not guarantee accurate | reproducible bias

Every laboratory can reproduce the same biased result if they share the same flawed reference or method. Reproducibility is consistency across conditions, not closeness to truth.

19. Reliability is broader consistency/stability language | reliability

Reliability has domain-specific meanings across psychometrics, engineering and forensic science. Broadly it concerns dependable/consistent results, but the exact coefficient or condition must be named in technical writing.

20. Reliable can still be invalid | reliability vs validity

A questionnaire can consistently measure the wrong construct. High reliability does not prove validity.

21. Validity asks whether we measure what we intend | validity

Validity is broader than accuracy in construct measurement. In education/psychology, validity concerns the interpretation/use of scores and whether evidence supports the intended construct/inference.

22. Construct validity is not calibration | construct validity

For abstract constructs such as motivation, there may be no physical reference standard. Validity evidence replaces simple calibration-to-truth logic.

23. Calibration compares against a standard/reference | calibration

NIST describes calibration as comparison of a device under test with an established standard. Calibration can estimate offset and uncertainty and support correction.

24. Calibration is not adjustment automatically | calibration vs adjustment

Calibration determines relation/offset to a standard; adjustment changes the instrument or result. A device can be calibrated and found biased without immediately being adjusted.

25. Correction compensates known systematic effect | correction

A correction factor/value compensates estimated systematic error. Because the error is not known perfectly, correction does not eliminate uncertainty.

26. Uncertainty quantifies doubt around a measurement result | uncertainty

Measurement uncertainty characterises the dispersion of values that could reasonably be attributed to the measurand under the framework. It is not the same as the unknowable exact measurement error.

27. Error and uncertainty are not synonyms | error vs uncertainty

Error is the difference from a reference/true value; uncertainty quantifies doubt about the measurement result/reference relationship. We may estimate uncertainty without knowing exact error.

28. Standard uncertainty is a technical quantity | standard uncertainty

Metrology expresses uncertainty using standard uncertainties and combined/expanded forms. General learners should recognise the vocabulary and rely on field-specific methods for calculation.

29. Expanded uncertainty uses a coverage factor | expanded uncertainty

Expanded uncertainty multiplies combined standard uncertainty by a coverage factor under a stated convention. “±X” should not be interpreted without knowing what X represents.

30. Resolution is the smallest meaningful display/increment | resolution

An instrument can display 0.001 units without being accurate to 0.001. Resolution is not precision or accuracy.

31. Sensitivity is response to change, not accuracy | sensitivity

In measurement systems, sensitivity concerns how output changes with input. In diagnostic testing, sensitivity has another definition. It is not a synonym for precision.

32. Specificity is domain-specific diagnostic classification | specificity

Diagnostic sensitivity/specificity concern classification performance, not metrological accuracy in the simple target-board sense. Keep the domain definition.

33. Stability asks whether measurement behaviour changes over time | stability

A device can be accurate today but drift over months. Stability concerns long-term change in bias/measurement characteristics.

34. Drift is systematic change over time | drift

Instrument drift changes calibration/bias gradually. Drift belongs to temporal measurement behaviour, not random scatter.

35. Agreement differs from correlation | agreement

Two instruments can correlate perfectly while one reads 10% higher than the other. Correlation tracks co-movement; agreement asks closeness of values.

36. Concordance is agreement-oriented | concordance

Concordance measures aim to capture agreement rather than only association. Use the specific method when technical interpretation matters.

37. Inter-rater reliability is scorer consistency | inter-rater

Two markers can agree consistently on scores but both may be biased relative to an external standard. Reliability and validity remain distinct.

38. Intra-rater reliability is self-consistency | intra-rater

A marker may score the same script similarly on repeated occasions. That shows consistency, not correctness of the scoring rubric.

39. Test–retest reliability concerns stability across time | test-retest

If the construct itself is stable, similar results across occasions can support test–retest reliability. Real change in the construct can reduce apparent stability without making the test poor.

40. Internal consistency is not one-dimensional validity | internal consistency

Items can correlate highly because they are repetitive. High internal consistency does not prove the scale measures the intended construct well.

41. Measurement quality has several axes | multi-axis

A complete description can include accuracy, bias, precision, repeatability, reproducibility, stability and uncertainty. One adjective rarely summarises them all.

42. Data quality is broader than measurement accuracy | data quality

Completeness, timeliness, consistency, validity and lineage can matter even when numerical measurements are accurate. “High-quality data” is broader than “accurate measurements.”

43. Model accuracy is task-specific | model accuracy

In machine learning, “accuracy” can also mean a classification metric: correct predictions divided by all predictions. That is not the same definition as metrological accuracy.

44. Forecast accuracy is realised prediction performance | forecast accuracy

Forecast accuracy compares predictions with realised outcomes using an error metric. It is distinct from calibration of probabilistic forecasts and from measurement accuracy of input data.

45. Part I checkpoint: consistent, correct and valid are three different questions | 第一部分检查点

Ask: Do repeated results agree? Do they agree with a reference? Does the measurement represent the intended construct? These correspond roughly to precision/reliability, accuracy/trueness and validity. Never let one answer stand for all three.

Part II — Measurement quality across real domains | 第二部分:真实领域中的测量质量

46. Laboratory measurements can be precise but biased | 实验室

A balance reads every standard 0.20 g high with very little scatter. Repeatability is excellent, bias is positive, and uncorrected accuracy is poor.

47. Reference-material quality limits calibration claims | reference

A calibration is only as trustworthy as its reference chain and uncertainty. Agreement with an uncertain reference cannot create certainty beyond that reference.

48. Inter-laboratory comparison tests reproducibility | inter-lab

Different laboratories apply a method to comparable material. If results agree under changed operators, equipment and locations, reproducibility is supported.

49. Same-lab repeated runs test repeatability | same lab

Repeating a method on the same instrument under similar short-term conditions primarily tests repeatability, not broad reproducibility.

50. Manufacturing: process precision differs from centering | manufacturing

A machining process can produce parts with tiny spread around the wrong diameter. Precision is high; process mean is off target. Quality needs both centering and spread relative to tolerance.

51. Calibration drift can make yesterday’s accuracy disappear | drift

An instrument calibrated last month can drift over time. Calibration history and stability determine whether the earlier accuracy statement still applies.

52. Gauge repeatability and reproducibility separates sources | Gage R&R

Manufacturing measurement-system studies often separate within-operator/equipment repeatability from between-operator or changed-condition reproducibility. The exact method is domain-specific.

53. Tolerance is not measurement uncertainty | tolerance vs uncertainty

A product tolerance defines acceptable product values. Measurement uncertainty describes doubt in the measured value. A part near the limit can be hard to classify if uncertainty is large.

54. Specification is not instrument resolution | specification

A specification may require ±0.1 mm, while the instrument displays 0.01 mm. Display resolution alone does not prove the instrument can measure accurately enough for the specification.

55. Education: marker agreement is reliability, not validity | scoring

Two teachers can consistently assign the same score to essays. That supports inter-rater reliability. Whether the rubric captures writing quality as intended is a validity question.

56. Rubric consistency can coexist with systematic harshness | scoring bias

Two markers may agree with each other but both score systematically lower than an external benchmark. Agreement/reliability does not rule out shared bias.

57. Test–retest reliability depends on construct stability | education

A vocabulary test given twice should correlate closely if the learner’s underlying ability has not changed. Real learning between tests can reduce similarity without indicating bad measurement.

58. Internal consistency can be inflated by redundant items | item redundancy

Asking nearly the same question ten times can create high internal consistency while narrowing construct coverage. Reliability can rise while validity worsens.

59. Survey measurement error includes respondent and wording effects | survey

Answers can vary because of question wording, interviewer effects, recall or social desirability. Sampling error and measurement error are different uncertainty sources.

60. Sample precision is not measurement accuracy | survey precision

A huge sample can produce a very narrow sampling interval around a biased survey answer. More respondents reduce some random sampling uncertainty but do not automatically fix wording or nonresponse bias.

61. Nonresponse bias is systematic representativeness error | nonresponse

If people who respond differ systematically from nonrespondents, a precise estimate can still be biased. Precision of the sample statistic is not population accuracy.

62. Census does not eliminate measurement error | census

Measuring every unit removes sampling error from sampling but not response, coverage, classification or processing errors.

63. Sensor noise is random-like scatter | sensors

Electronic sensors can fluctuate around the true signal due to noise. Repeated averaging can improve precision if noise is sufficiently independent, but drift/offset remain.

64. Offset is systematic difference | offset

An offset shifts readings by a nearly constant amount. Calibration/correction can reduce known offset while leaving residual uncertainty.

65. Gain error changes slope, not only offset | scale error

An instrument can be correct at zero but increasingly wrong as values rise. Calibration may need both offset and scale/sensitivity correction.

66. Nonlinearity is another calibration defect | nonlinear response

A sensor’s response may not maintain the expected relation across its range. One calibration point cannot establish accuracy everywhere.

67. Hysteresis makes reading depend on direction/history | hysteresis

Some instruments produce different readings for the same input depending on whether the input was approached from above or below. Repeatability studies must define operating conditions.

68. Environmental conditions affect reproducibility | environment

Temperature, humidity, vibration or operator conditions can change results. A method reproducible only under tightly controlled conditions may fail in the field.

69. Human judgment adds rater variability | human measurement

Clinical ratings, essay scores and inspections involve human interpretation. Training can improve agreement but may also spread a shared bias if the standard itself is flawed.

70. Agreement must be assessed on the correct scale | agreement scale

Correlation, percent agreement, kappa, limits of agreement and intraclass correlation answer different questions. Use the method suited to measurement type and decision.

71. Machine learning classification accuracy is a proportion | ML accuracy

Classification accuracy is correct predictions divided by all predictions. It can be misleading with severe class imbalance and is not the same concept as metrological closeness to a reference value.

72. Precision in machine learning is a different metric | ML precision

In classification, precision means the share of predicted positives that are truly positive. This is a domain-specific term unrelated to metrological repeatability/dispersion.

73. Recall is another classification denominator | recall

Recall/sensitivity asks what share of actual positives were detected. Precision and recall use different denominators; do not treat them as general “accuracy.”

74. Calibration in probabilistic prediction has another meaning | probabilistic calibration

A probability model is calibrated when stated probabilities align with outcome frequencies across comparable cases. This is different from instrument calibration against a physical standard.

75. Discrimination differs from calibration | model quality

A model can rank high-risk cases well but assign probabilities that are systematically too high or low. Discrimination and calibration are separate model-quality axes.

76. Forecast error metrics encode different penalties | forecasting

MAE, RMSE, percentage errors and other metrics weight mistakes differently. “More accurate forecast” should ideally name the metric/horizon.

77. Mean error can hide cancellation | bias in forecasts

Positive and negative forecast errors can average near zero even when absolute errors are large. Mean error tracks directional bias, not total accuracy.

78. MAE measures average absolute magnitude of error | MAE

Mean absolute error ignores sign and summarises typical absolute discrepancy. It is on the outcome’s unit scale.

79. RMSE penalises larger errors more | RMSE

Root mean squared error squares discrepancies before averaging, giving large misses more influence. MAE and RMSE can rank models differently.

80. Percentage error can behave badly near zero | relative error

Relative/percentage error can explode when actual values are near zero. Metric choice should match the scale and decision.

81. Accuracy across subgroups can differ | subgroup performance

An overall measurement/model metric can hide systematic errors for subgroups. Check calibration, bias and reliability across relevant populations rather than only globally.

82. Measurement invariance concerns comparability across groups | invariance

In psychometrics, measurement invariance asks whether a construct is measured comparably across groups. Equal reliability alone does not establish comparable meaning.

83. Gold standard is rarely literally perfect | gold standard

A “gold standard” is a best available reference, not necessarily error-free truth. Reference uncertainty should remain part of interpretation.

84. Ground truth can be constructed or uncertain | ground truth

Machine learning and remote sensing use ground truth, but labels can contain human or measurement error. The term should not hide uncertainty in the reference.

85. Benchmark can be imperfect | benchmark

A benchmark may be a reference model or standard rather than truth. Beating a benchmark does not prove absolute accuracy.

86. Reproducible code is not automatically reproducible science | computational reproducibility

Running the same code on the same data can reproduce a result. Independent data, measurements or methods may be needed for broader scientific reproducibility/replicability.

87. Replication asks whether findings recur in new studies | replication

Research communities distinguish reproducibility and replication in different ways. State whether you mean same data/code or independent data/design.

88. Versioning is part of reproducibility | version control

Software, data and parameter versions can change results. A reproducible workflow records the computational environment and provenance.

89. Traceability links measurements to references | traceability

Metrological traceability connects a result to a reference through a documented unbroken calibration chain with stated uncertainties. It is stronger than saying “we calibrated the device once.”

90. Part II checkpoint: quality depends on the question and conditions | 第二部分检查点

Ask whether you care about closeness to reference, mutual agreement, same-condition repeatability, changed-condition reproducibility, long-term stability, construct validity or prediction performance. One word—“accurate”—cannot carry every quality dimension.

Part III — Mandarin-to-English measurement-quality control | 第三部分:中文母语学习者的测量质量词汇转换

91. 准确 = accurate / correct, depending job | 准确

“测量准确” → the measurement is accurate if closeness to a reference is established. “答案正确” → the answer is correct. Do not use accurate for every kind of correctness.

92. 准确度 = accuracy | 准确度

In metrology, accuracy is closeness to accepted reference. In machine learning, accuracy can be a classification proportion. Domain changes definition.

93. 精确 can mean precise or accurate in ordinary Chinese | 精确

Chinese 精确 often blends “exact” and “precise.” In technical English, decide whether the point is repeated agreement (precise) or closeness to reference (accurate).

94. 精度 is ambiguous across engineering contexts | 精度

Depending source, 精度 may mean accuracy, precision, resolution or machining tolerance. Translate from the specification, not the Chinese label alone.

95. 精密度 = precision | 精密度

When the technical meaning is closeness among repeated results, use precision. Do not translate it as accuracy merely because the readings look “good.”

96. 真度 / 正确度 can map to trueness | trueness

Standards/metrology may use a Chinese term corresponding to trueness, meaning closeness of the average to reference. Verify the standard vocabulary because translations vary.

97. 误差 = error | 误差

“测量误差” → measurement error. Error is a discrepancy relative to reference/true value, not the same as uncertainty.

98. 随机误差 = random error | 随机误差

Use for unpredictable scatter components across repeated measurements. It contributes to imprecision.

99. 系统误差 = systematic error | 系统误差

Use for persistent directional components of measurement error. Repetition alone does not remove it.

100. 偏差 can be bias, deviation or difference | 偏差

Systematic measurement 偏差 → bias. A single value’s departure from mean → deviation. General difference → difference/discrepancy. Chinese 偏差 is highly context-sensitive.

101. 偏置 = bias | 偏置

In statistics/machine learning, 偏置 commonly maps to bias. It may mean systematic estimation error or model bias depending framework.

102. 偏离 = deviation / departure | 偏离

“偏离参考值” → deviates from the reference value. This is descriptive and need not imply systematic bias across repeated measurements.

103. 可靠性 = reliability | 可靠性

In measurement/testing, reliability concerns consistency/dependability under a specified framework. In engineering, reliability can also mean probability a system performs without failure. Always identify domain.

104. 稳定性 = stability | 稳定性

For measurement systems, stability concerns change in measurement behaviour over time. It is not the same as repeatability over a short period.

105. 重复性 = repeatability | 重复性

Use when repeated results are obtained under essentially the same procedure, operator, equipment, place and short time conditions.

106. 再现性 = reproducibility | 再现性

Use when agreement is assessed under changed conditions such as different operators, labs or instruments. State which conditions changed.

107. 可重复 = repeatable / reproducible, depending Chinese usage | 可重复

Chinese 可重复 can be ambiguous. Same-team/same-condition repetition → repeatable. Independent/changed-condition recreation may be reproducible. Research communities also use terminology differently; define your sense.

108. 可复现 = reproducible / replicable, depending research convention | 可复现

Computational research may use reproducible for same data/code; other fields distinguish replication with new data. State the convention instead of assuming one universal mapping.

109. 校准 = calibration | 校准

Use calibration for comparison of a measurement system against a reference/standard and characterisation of its response. Calibration does not automatically mean the device was adjusted.

110. 标定 can be calibration / characterization | 标定

Engineering Chinese may use 标定 for calibration or determination of scale/response. Follow the instrument standard and documentation.

111. 修正 = correction / adjustment | 修正

A known offset may be corrected mathematically; an instrument may be adjusted physically. Do not call every correction a calibration.

112. 不确定度 = measurement uncertainty | 不确定度

In metrology, uncertainty is a quantified property of doubt around the measurement result. It is not the same as known error.

113. 分辨率 = resolution | 分辨率

Resolution concerns the smallest meaningful/displayed increment. High resolution does not guarantee high accuracy or precision.

114. 灵敏度 = sensitivity, but domain changes meaning | 灵敏度

Instrument sensitivity describes response per input change; diagnostic sensitivity describes true-positive detection. Do not merge these definitions.

115. 特异度 = specificity | 特异度

In diagnostic classification, specificity concerns correct identification of negatives. It is not general measurement accuracy.

116. 有效性 = validity | 有效性

For tests/constructs, validity concerns whether evidence supports the intended interpretation/use. “有效” is not simply “accurate.”

117. 一致性 = consistency / agreement | 一致性

Results being similar can indicate consistency/agreement. It does not automatically establish reliability coefficient, accuracy or validity.

118. 符合度 = agreement / concordance / goodness of fit | 符合度

Chinese 符合度 is context-sensitive: method agreement, model fit or compliance with a standard may require different English terms.

119. 漂移 = drift | 漂移

Measurement drift is gradual systematic change over time. It can degrade accuracy even if short-term repeatability remains strong.

120. 参考值 = reference value | 参考值

Use reference value when comparing measurement accuracy. Benchmark may be appropriate for performance comparison but is not necessarily a metrological reference.

121. 真值 = true value, but often unknowable | 真值

Metrology recognises that the true value may be unknowable; accepted/reference values are used operationally. Avoid pretending the reference is perfect truth.

122. 实测值 = measured / observed value | 实测值

A measured value is the result obtained from the process. It can be precise, biased, uncertain or inaccurate relative to the reference.

123. 偏高 / 偏低 = positively / negatively biased or systematically high/low | 偏高偏低

If repeated measurements are systematically above reference, say positively biased/systematically high. One single high reading is merely above the reference, not evidence of bias.

124. 稳定但不准 = precise/repeatable but inaccurate | 稳定但不准

“结果很稳定,但总是低 0.2” → The measurements are highly repeatable/precise but show a consistent negative bias.

125. Part III checkpoint: 准确、精度、偏差 and 稳定 cannot share one English word | 第三部分检查点

Ask: closeness to reference? mutual agreement? systematic offset? same-condition consistency? changed-condition consistency? long-term drift? construct validity? The measurement-quality job determines the English term.

Part IV — Measurement-quality failure laboratory | 第四部分:测量质量失误实验室

126. Precise but biased | 精密但有偏

Reference = 10.00. Repeated readings = 9.80, 9.81, 9.80, 9.79. Scatter is tiny; systematic offset is about -0.20. The system is precise/repeatable but inaccurate before correction.

127. Accurate on average but imprecise | 平均准但离散大

Readings = 8, 12, 9, 11 around reference 10. Mean is near reference but individual measurements scatter. Good average trueness does not make each reading precise.

128. More repeats reduce noise but preserve bias | 多测几次也会稳定地错

A sensor is always 2 units high with random noise ±0.5. Averaging 1,000 readings gives a very stable estimate near +2 bias. Precision improves; systematic error remains.

129. Perfect correlation but poor agreement | 高相关但不一致

Instrument B always reports twice Instrument A. Correlation can be +1 while values disagree completely in scale. Association is not agreement.

130. Strong agreement around the wrong reference | shared bias

Two laboratories use the same miscalibrated standard and reproduce the same biased result. Reproducibility is good; accuracy relative to the correct reference is poor.

131. Calibration certificate does not guarantee current accuracy | drift

An instrument was calibrated six months ago but has since drifted. A historical calibration is evidence about a past state plus traceability, not proof of present performance.

132. High display resolution creates fake confidence | digits

A device displays 12.34567 but its uncertainty is ±0.2. Extra digits are resolution/display precision, not justified measurement accuracy.

133. Tight tolerance with weak measurement system | decision risk

A product must be within ±0.1 while measurement uncertainty is ±0.15. The instrument cannot confidently decide near the specification boundary.

134. Reliable test measures the wrong construct | reliability without validity

A vocabulary test consistently measures spelling speed when the intended construct is lexical depth. Scores are stable but the interpretation is invalid.

135. Valid-looking content but low reliability | unstable measure

Items appear relevant to critical thinking but scoring varies wildly across occasions/raters. Construct coverage alone does not create dependable measurement.

136. Inter-rater agreement hides shared rubric bias | markers

Two markers agree almost perfectly because both interpret a flawed rubric identically. Reliability is strong; validity relative to intended writing quality may be weak.

137. Large sample creates precise biased estimate | survey

A million respondents answer a leading question. Sampling uncertainty is tiny, but measurement/response bias can remain large. Precision is not truth.

138. Census still contains measurement error | census

Counting everyone eliminates sampling-from-population error but not misclassification, missing units, duplicate records or respondent error.

139. High ML classification accuracy from class imbalance | model metric

If 99% of cases are negative, always predicting negative yields 99% classification accuracy but detects no positives. Domain-specific accuracy metrics need denominator context.

140. ML precision is not metrology precision | false friend

Classification precision is positive predictive value among predicted positives; metrological precision is repeat-measurement agreement. Same English word, different technical owner.

141. Probabilistic model discriminates well but is miscalibrated | model quality

A model ranks risky cases correctly but predicts 90% probability where observed frequency is 60%. Discrimination can be high while calibration is poor.

142. Mean forecast error near zero hides large misses | cancellation

Errors +20 and -20 average to zero. Directional bias is zero, but accuracy is poor. Use absolute/squared error metrics for magnitude.

143. RMSE worse than MAE because of one extreme miss | metric sensitivity

One large forecast miss affects RMSE strongly because errors are squared. Different accuracy metrics answer different penalty questions.

144. Reproducible code reproduces the same bug | computational reproducibility

Everyone running the same flawed code gets the same result. Computational reproducibility is high; validity/correctness may be low.

145. Independent replication fails | scientific replication

A published effect can be perfectly reproducible from its original files but fail in independent new data. Reproduction of computation and replication of phenomenon are distinct.

146. “Gold standard” has its own error | reference uncertainty

A new method is compared with an imperfect reference. Apparent disagreements can partly come from the reference. Accuracy claims inherit reference uncertainty.

147. Correction removes estimated bias but leaves uncertainty | correction

Subtracting a known calibration offset improves trueness, but the offset estimate itself is uncertain. Corrected does not mean exact.

148. Part IV checkpoint: “consistent” and “correct” are different axes | 第四部分检查点

The recurring failure is to treat one quality dimension as all quality: correlation as agreement, precision as accuracy, repeatability as reproducibility, reliability as validity, calibration as certainty, or narrow intervals as truth. Name the axis you actually measured.

149. Build the measurement FENCE | 测量质量 FENCE

FenceQuestionJob
F0 MeasurandWhat exactly is measured?define quantity/construct
F1 ReferenceCompared with what?accuracy/trueness
F2 ErrorHow far from reference?discrepancy
F3 BiasIs there systematic offset?directional error
F4 PrecisionHow tightly do repeats agree?random scatter
F5 ConditionsSame or changed conditions?repeatability/reproducibility
F6 StabilityDoes behaviour change over time?drift
F7 ValidityDoes it measure the intended construct?interpretation
F8 ClaimWhat quality word is justified?accuracy/precision/reliability/etc.

150. Build a measurement-control card | 测量控制卡

FieldExample
Measurandmass
Referencecertified 100.00 g standard
Mean reading99.80 g
Bias-0.20 g
Repeat spreadsmall
Repeatabilityhigh
Reproducibilitynot yet tested across labs
Uncertaintyreported separately

151. Practice A — accurate, precise, both or neither? | 练习 A

  1. Readings tightly around the reference.
  2. Readings tightly far above the reference.
  3. Readings widely scattered around the reference.
  4. Readings widely scattered far from the reference.

Classify the four patterns and explain bias versus random error.

152. Practice B — repeatability or reproducibility? | 练习 B

  1. Same operator, same device, ten repeats in ten minutes.
  2. Five laboratories use the same method on the same reference sample.
  3. Same laboratory repeats six months later with a new operator.
  4. Different instrument models in different locations.

Name the condition and specify what changed.

153. Practice C — correlation or agreement? | 练习 C

Method B always equals 1.2 × Method A. Correlation is nearly perfect. Explain why this is insufficient evidence that the methods agree interchangeably.

154. Practice D — reliability or validity? | 练习 D

  1. A test gives similar scores twice.
  2. Two markers give similar scores.
  3. A test accurately represents the intended construct.
  4. Items are very similar to one another but cover only one narrow skill.

155. Practice E — Mandarin repair | 练习 E

  1. 仪器读数非常稳定,但比参考值低 0.2。
  2. 重复性很好,但跨实验室再现性尚未验证。
  3. 校准后仍然存在测量不确定度。
  4. 两种方法相关很高,但系统性偏差导致它们不能互换。
  5. 测试可靠性高,不代表效度一定高。

156. A 15-minute measurement-language session | 15 分钟训练

TimeTask
0–3Name measurand and reference.
3–6Inspect bias and spread.
6–9Identify repeatability conditions.
9–12Ask what changed for reproducibility.
12–15Write one accurate quality statement.

157. Seven-day mastery route | 七天路线

DayFocus
1accuracy / precision / trueness
2bias / error / random vs systematic
3repeatability / reproducibility
4calibration / uncertainty / traceability
5reliability / validity / agreement
6Mandarin transfer
7full FENCE audit

158. First weak-link diagnostic | 弱点诊断

SymptomWeaknessRepair
I call consistent readings accurate.precision/accuracytarget-board cases
I average away bias.random/systematicrepetition examples
I mix repeatability and reproducibility.conditionssame-vs-changed drills
I call calibration adjustment.process distinctionreference/correction sequence
I think reliable means valid.construct reasoningwrong-construct examples
I translate 精度 as accuracy every time.bilingual ambiguityaccuracy/precision/resolution contrasts

159. Reading harvest | 阅读提取

Collect full bundles: accepted reference value, systematic bias, repeatability conditions, reproducibility across laboratories, calibrated against, expanded uncertainty, inter-rater reliability, construct validity. The condition/reference phrase carries the meaning.

160. Writing activation | 写作激活

Use: measurand → reference → bias → repeat spread → conditions → uncertainty → quality conclusion. Avoid one-word claims such as “highly accurate” without evidence.

161. Speaking activation | 口语激活

Explain one measurement system in 60 seconds. The listener should know whether it is close to reference, repeatable, reproducible, stable and valid for the intended use.

162. Paraphrase test: preserve the quality axis | 改述测试

Original: “The method showed excellent repeatability but a negative bias.” Unsafe: “The method was highly accurate.” Safe: “Repeated measurements were tightly consistent, but their average was systematically below the reference.”

163. Summary test: do not erase the reference | 摘要测试

A result can only be described as accurate relative to a reference or validated outcome. If the reference is uncertain, preserve that limitation.

164. Canonical ownership boundary | 本课所有权边界

Lesson 044 owns measurement-quality language: accuracy, precision, trueness, bias, error, repeatability, reproducibility, reliability, validity, calibration, uncertainty, stability and agreement. Lesson 039 owns distributional spread; Lesson 041 association/correlation; Lesson 043 future-claim and forecast language.

165. Recommended reference floor | 推荐参考资源

166. SEO language map | 本课关键词范围

accuracy vs precision, bias vs error, random vs systematic error, repeatability vs reproducibility, reliability vs validity, calibration vs accuracy, measurement uncertainty, correlation vs agreement, precise but inaccurate, accurate but imprecise, 准确度英语, 精密度英语, 偏差误差英语, 重复性再现性英语, 可靠性效度英语.

167. Your assignment | 本课作业

Choose one real measurement process. Record at least five repeated measurements of a stable reference, then write a 300-word quality report distinguishing reference, bias, spread, repeatability, uncertainty and what evidence would still be needed for reproducibility.

168. Why this owner earns longform depth | 为什么需要长文深度?

The reader job is not memorising “accuracy vs precision.” Real measurement quality includes reference uncertainty, systematic bias, random scatter, operator/lab changes, drift, construct validity, calibration, forecast/model metrics and domain-specific false friends. A consistent number can be wrong in several different ways.

169. Final rule one | 最后一条规则一

Consistency tells you whether measurements agree with each other. Accuracy asks whether they agree with the reference.

一致性告诉你测量彼此是否相符;准确度问它们是否接近参考值。

170. Final rule two | 最后一条规则二

A precise instrument can be consistently wrong.

一个很精密的仪器,也可以稳定地测错。