Series ID: EDKS-BASIC-ZH-0086 · Lesson No.086 · C1→C2-Direction Evidence Literacy
Reading Complex Evidence and Method Claims | 阅读复杂证据与方法主张
Advanced reading asks not only “What did the study find?” but also “What was actually measured, compared, controlled and justified?”
At high levels, readers must separate the wording of a conclusion from the strength of the method behind it. A confident sentence can still rest on limited evidence; a cautious sentence can reflect strong scientific discipline.
高级证据阅读的关键不是记住很多研究术语,而是判断:这个方法到底允许作者说到哪一步?
Chinese Edition Hub · 中文版入口 · ← Lesson No.085 · Advanced Collocation and Semantic Precision
The evidence-reading map | 证据阅读地图
- research question
- sample and population
- measurement
- comparison group
- randomisation
- confounding
- effect size
- uncertainty
- correlation vs causation
- replication
- external validity
- claim calibration
1. Start with the research question | 先看研究问题
Ask what the study was designed to answer.
2. Descriptive question | 描述性问题
How many participants preferred Option A?
This can be answered by appropriate observation or survey data.
3. Comparative question | 比较问题
Did Group A perform differently from Group B?
This requires a meaningful comparison.
4. Causal question | 因果问题
Did the intervention cause the difference?
This demands much stronger design control.
5. Sample | 样本
The sample is the group actually studied.
6. Population | 总体
The population is the broader group the authors may wish to understand.
7. Sample ≠ population | 样本不等于总体
A study of 80 volunteers from one school cannot automatically represent all learners.
8. Sampling method | 抽样方法
- random sample
- convenience sample
- volunteer sample
- purposive sample
9. Convenience sample | 便利样本
Easy to recruit, but may differ systematically from the wider population.
10. Volunteer bias | 志愿者偏差
People who volunteer may be more motivated or interested than those who do not.
11. Sample size | 样本量
Larger samples can reduce some forms of random uncertainty, but size alone does not repair poor sampling or poor measurement.
12. Bigger is not automatically better | 大不一定好
A huge biased sample can still produce a misleading estimate.
13. Measurement | 测量
Ask what the researchers actually measured.
14. Construct | 构念
Abstract ideas such as:
- confidence
- motivation
- learning
- stress
- productivity
must be turned into measurable indicators.
15. Operationalisation | 操作化
How was the concept measured in practice?
For example, “learning” might mean:
- test score
- retention after one week
- task accuracy
- self-reported improvement
16. Different measures answer different questions | 不同测量回答不同问题
Self-reported confidence is not the same as measured performance.
17. Reliability | 信度
Would the measure produce reasonably consistent results under similar conditions?
18. Validity | 效度
Does the measure actually capture the concept the claim is about?
19. Face validity is not enough | 看起来合理不够
A questionnaire can look sensible while still measuring the wrong construct.
20. Comparison group | 对照/比较组
Without a comparison, improvement over time may have many explanations.
21. Before–after design | 前后设计
Useful for change, but weak for causality if many things changed at the same time.
22. Control group | 对照组
A control group can help estimate what might have happened without the intervention.
23. Random assignment | 随机分配
Random assignment can reduce systematic pre-existing differences between groups in well-designed experiments.
24. Random sampling vs random assignment | 随机抽样 vs 随机分配
- random sampling → helps represent a population
- random assignment → helps causal comparison
They solve different problems.
25. Confounding | 混杂因素
A confound is another factor related to both the supposed cause and the outcome.
26. Example of confounding | 混杂例子
If students who choose an optional course also study more outside class, higher scores may not be caused by the course alone.
27. Alternative explanation | 替代解释
Strong reading asks what else could produce the same pattern.
28. Correlation | 相关
Two variables move together.
29. Correlation does not by itself establish causation | 相关不等于因果
Possible explanations include:
- X causes Y
- Y causes X
- Z causes both
- chance
30. Association language | 相关语言
- associated with
- linked to
- correlated with
These should not silently become caused.
31. Causal language | 因果语言
- caused
- led to
- resulted in
- produced
These require stronger support.
32. Effect size | 效应大小
A statistically detectable difference can still be very small in practical terms.
33. Statistical significance vs practical importance | 统计显著 vs 实际重要
These are not the same question.
34. Absolute change | 绝对变化
From 2% to 3% = 1 percentage-point increase.
35. Relative change | 相对变化
From 2% to 3% = 50% relative increase.
Both are mathematically correct; the frame affects interpretation.
36. Denominator | 分母
“100 cases” is hard to interpret without knowing out of how many.
37. Confidence interval | 置信区间
A confidence interval expresses uncertainty around an estimate under a statistical model.
At this level, the key reading skill is to notice whether the interval is wide or narrow and what values remain plausible.
38. Uncertainty is information | 不确定性也是信息
A careful study often reports uncertainty instead of pretending to know an exact value.
39. Precision of measurement | 测量精度
Do not report more decimal places than the measurement process justifies.
40. Missing data | 缺失数据
Who dropped out or failed to respond?
41. Attrition | 流失
If many participants leave, the remaining group may no longer represent the original sample.
42. Survivorship | 幸存者偏差
Studying only successful completers can hide failure patterns.
43. Blinding | 盲法
In some study types, keeping participants or assessors unaware of conditions can reduce expectation effects.
44. Observer effects | 观察者效应
Measurement itself can influence behaviour.
45. Self-report limits | 自我报告限制
People may report beliefs, memories or preferences imperfectly.
46. Objective measurement is not automatically perfect | “客观”测量也有局限
A digital metric may be precise while measuring the wrong thing.
47. Replication | 重复研究
One study is less informative than a pattern reproduced across appropriate settings.
48. Replication is not simple duplication | 重复不等于照抄
Useful replication may test whether the result survives different populations, researchers or contexts.
49. External validity | 外部效度
Can the finding generalise beyond the study setting?
50. Internal validity | 内部效度
How well does the design support the causal explanation within the study?
51. Internal vs external trade-off | 内外效度取舍
A tightly controlled study may establish a mechanism well but apply imperfectly to real-world settings.
52. Ecological validity | 真实情境适用性
Does the task resemble the environment where the conclusion will be used?
53. Method claim vs result claim | 方法主张 vs 结果主张
Result:
Group A scored higher.
Method claim:
The design allows us to attribute the difference to the intervention.
The second requires much more justification.
54. Headline vs study | 标题 vs 研究
Headlines may compress “is associated with” into “causes”. Always inspect the underlying claim where possible.
55. Meta-analysis | 荟萃分析
A meta-analysis combines results across studies using statistical methods.
But its strength depends on the included studies, comparability and methods.
56. Systematic review | 系统综述
A systematic review uses explicit methods to identify and evaluate relevant studies.
57. Expert review vs empirical evidence | 专家评论 vs 实证证据
Expert interpretation can be valuable, but it is not the same evidence type as direct measurement.
58. Pre-registration / planned analysis | 预先计划
Pre-specifying hypotheses and analyses can reduce some forms of hindsight-driven interpretation.
59. Multiple comparisons | 多重比较
When many tests are run, some differences may appear by chance.
60. Selective reporting | 选择性报告
Ask whether only successful outcomes are highlighted.
61. Negative result | 阴性结果
“No statistically clear difference” does not automatically mean “the two treatments are identical”.
62. Absence of evidence vs evidence of absence | 缺少证据 vs 证据表明不存在
These are different conclusions.
63. Bayesian-style reading intuition | 贝叶斯式直觉
New evidence should update confidence rather than force every question into “proved / disproved”.
64. Method vocabulary should not intimidate | 方法词汇不是为了吓人
The core questions remain simple:
- Who was studied?
- What was measured?
- Compared with what?
- What else could explain it?
- How uncertain is the estimate?
- How far can we generalise?
65. Mandarin transfer: “证明”常被过度翻成 prove | 注意强度
Many empirical studies support, suggest or indicate rather than prove in the strongest sense.
66. Mandarin transfer: “有效”必须先定义 | effective for what?
Ask what outcome, time scale and population are meant.
67. Reading practice A | 阅读 A
A voluntary survey of 300 course completers found high satisfaction. The report concluded that the programme was highly effective.
Identify at least three method questions before accepting the conclusion.
68. Reading practice B | 阅读 B
After a new schedule was introduced, productivity rose by 8%. During the same period, the company also reduced meetings and adopted new software.
Why is a simple causal claim unsafe?
69. Reading practice C | 阅读 C
The treatment group improved by 3 points while the comparison group improved by 2 points.
Which difference matters for estimating the treatment effect?
70. Claim-calibration practice | 主张校准
Weak design → cautious wording.
Strong design → stronger wording may be justified.
71. Error clinic | 常见问题
| Problem | Repair |
|---|---|
| Large sample = good study. | Check sampling and measurement. |
| After X, therefore because of X. | Check confounding/comparison. |
| Statistically significant = important. | Check effect size. |
| No significance = no effect. | Check uncertainty and power. |
| Survey preference = performance. | Match measure to claim. |
72. First weak link diagnosis | 第一个卡点诊断
- I accept result wording → method/claim separation.
- I confuse sample/population → generalisation.
- I infer causality from timing → confounding.
- I cannot interpret percentages → absolute/relative framing.
- I overreact to statistical terms → evidence literacy.
73. Seven-day training cycle | 七天训练
| Day 1 | sample/population | 样本总体 |
| Day 2 | measurement/validity | 测量效度 |
| Day 3 | comparison/confounding | 对照混杂 |
| Day 4 | correlation/causation | 相关因果 |
| Day 5 | effect size/uncertainty | 效应不确定 |
| Day 6 | generalisation/replication | 外推重复 |
| Day 7 | three-study method audit | 综合 |
74. Self-test | 自测
Distinguish sample from population, measurement from construct, correlation from causation, statistical significance from practical importance, and internal from external validity; calibrate conclusions to method strength.
75. For parents and teachers | 给家长和老师
Teach method reading through concrete questions before technical terminology.
Ask learners to rewrite overconfident conclusions into evidence-matched versions.
For Mandarin speakers, contrast 证明/prove with suggest/indicate/support carefully.
76. Final real-world challenge | 最终真实任务
- Read three study summaries.
- Identify sample and population.
- Identify the measured outcome.
- Check comparison groups.
- List plausible confounders.
- Separate correlation and causation.
- Compare absolute and relative changes.
- State uncertainty.
- Judge generalisability.
- Rewrite each conclusion at the strongest defensible level.
Next: Lesson No.087 | 下一课
The next lesson develops speaking across registers and audiences: preserving the same core meaning while adapting detail, directness, vocabulary, evidence and interaction style for different listeners.
Lesson No.087 · Speaking Across Registers and Audiences · 跨语体与受众口语表达
Reference floor: C1→C2-direction evidence reading; method terms are always tied back to what a design can and cannot support.