Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

How to Learn Advanced English (Chinese Edition) | Lesson No.020 | Write a Discussion That Explains Without Overclaiming | 第020课:把讨论写成解释,而不是把结果说得更大

Series ID: EDKS-ADV-ZH-0020 · How to Learn Advanced English (Chinese Edition) · Lesson No.020 · C1 → C2

Write a Discussion That Explains Without Overclaiming | 把讨论写成解释,而不是把结果说得更大

The Discussion section is where a research paper stops merely reporting what happened and begins deciding what the findings can reasonably mean. That makes it one of the most intellectually demanding parts of advanced academic English.

Discussion 是 research paper 里最容易“写得像学术”,却最容易在逻辑上失控的 section。Results 告诉读者发生了什么;Discussion 必须解释这些 findings 能支持什么、不能支持什么、怎样与既有研究连接、哪些 alternative explanations 仍然存在,以及 conclusion 最终应该有多强。

For Mandarin-speaking advanced learners, the difficulty is not only vocabulary. Chinese academic and professional discourse can tolerate forms such as “由此可见”, “这说明”, “因此证明”, “可见该方法有效”, or “这可能是由于” with different rhetorical force depending on context. Direct translation into English can silently increase certainty. This shows, this proves, therefore, because, must and should all carry inferential commitments. At C1–C2 level, you must control those commitments deliberately.

The operating principle of this lesson is simple:

finding → interpretation → test → boundary → implication.

先看 finding;再提出 interpretation;再让 interpretation 经受 alternative explanation、limitation 与 prior evidence 的测试;最后才决定 implication。


1. Why we are learning Discussion at this depth | 为什么这一课需要长篇深度训练

This Advanced English Chinese Edition exists to take Mandarin-speaking learners beyond “correct English” into independent C1–C2 performance. At this level, knowing a list of phrases such as the findings suggest that is not enough. A learner may write grammatically perfect English and still overclaim the evidence, hide a null result, treat a limitation as a ceremonial sentence, confuse explanation with proof, or recommend a policy that the study never tested.

That is why this lesson is not a phrase bank. It is a full operating system for Discussion. The goal is transfer: you should be able to read an unfamiliar paper, reconstruct the evidential state, distinguish what is observed from what is inferred, compare several plausible explanations, state the limitation that actually changes the claim, and write a final conclusion whose strength matches the evidence.

The same skill is useful outside research. Professional reports, strategic recommendations, policy briefs, evaluations, technical post-mortems and evidence-based proposals all require the same discipline: do not allow the story you want to tell to become stronger than the evidence you actually have.

2. The boundary with Results | Discussion 不是把 Results 再写一遍

Lesson 019 owns Results depth. There we learned to make the pattern visible without arguing ahead of the evidence. Discussion begins after that handoff.

Results might say:

The structured-feedback group scored 6.0 points higher on the one-week independent-revision task (95% CI 2.4–9.6).

Discussion might say:

The one-week advantage suggests that the structured condition supported some transfer beyond the practised text. However, because the intervention also required more explicit revision planning, the study does not isolate whether the benefit came from feedback structure itself or from the additional cognitive work imposed by the prompts.

The first sentence reports an estimate. The second paragraph begins interpretive work. It makes a claim about transfer, then immediately exposes a rival explanation. That is the level of control required in Discussion.

3. Manchester: Discussion is built from cycles | Discussion 不是一条直线

The University of Manchester Academic Phrasebank describes research Discussion sections as among the most complex parts of dissertations and articles. They often centre on important findings and repeat a series of moves around each finding: restatement, interpretation, comparison with prior literature, explanation, limitation, implication and future work.

University of Manchester Academic Phrasebank | Discussing Findings

This is important because many learners imagine Discussion as one large final argument. In practice, strong Discussions are often a sequence of smaller cycles:

Finding A → meaning → prior work → rival → limit.

Finding B → meaning → prior work → unexpected pattern → implication.

Finding C → null/negative result → uncertainty → future test.

The article’s overall conclusion emerges from the combined cycles. It should not be decided in advance and then forced onto each result.

4. Nature: interpretation must match the statistical evidence | Interpretation 不能比 evidence 更强

Nature Human Behaviour’s editorial guidance explicitly warns that Discussion sections can fail when interpretations do not match the statistical evidence, when limitations are missing, or when significance and implications are overstated. It also notes that transparent limitations increase credibility rather than weaken the work.

Nature Human Behaviour | Up Close and Personal

This gives us a strong Discussion rule:

The interpretation is not allowed to upgrade the evidence.

If Results show an association, Discussion cannot convert it into a demonstrated cause. If the confidence interval is wide, Discussion cannot write as if the magnitude were precisely known. If the primary outcome is null, Discussion cannot rescue the paper by promoting a positive secondary outcome without making that change of hierarchy explicit.

5. Nature Cancer: connect findings, knowledge, implications and caveats | Discussion 是 perspective section

A Nature Cancer editorial on scientific writing describes Discussion as the place where findings are put into perspective by connecting them with existing knowledge, implications, outstanding questions and future paths, while avoiding hype and acknowledging caveats and limitations.

Nature Cancer | The Craft (and Art) of Scientific Writing

This is a useful distinction: Discussion is not the place to make the Results sound more exciting. It is the place to make the Results more intelligible.

6. The first Discussion sentence should answer, not replay | 开头不要重新抄 Results

A weak Discussion often begins by reproducing numbers:

The intervention group scored 74.2 and the control group scored 68.1, a difference of 6.1 points.

The reader has just read that in Results.

A stronger opening answers the research question:

Structured feedback improved near-term independent revision in this sample, with the clearest advantage in evidence integration rather than sentence-level accuracy.

This sentence compresses the evidence into a finding-level answer. It is not yet a mechanism. It tells the reader what the study has learned.

7. Restatement is allowed; repetition is not | Restatement 有工作,repetition 没有

Restatement selects the result that matters for interpretation and expresses it at the right conceptual level. Repetition simply reproduces detail.

Results:

Mean revision scores differed by 6.0 points.

Discussion restatement:

The structured condition produced a clear near-term revision advantage.

Discussion overreach:

The structured condition transformed learners into stronger independent writers.

The first is data. The second is calibrated synthesis. The third expands the construct beyond what was measured.

8. Answer the primary question before celebrating secondary patterns | Main question first

If the primary outcome does not support the hypothesis, say so early. Do not begin with an attractive secondary finding.

Weak:

Learners reported greater confidence and valued the feedback highly.

when the primary writing outcome was null.

Stronger:

The intervention did not produce a clear improvement in the primary independent-writing outcome, although participants reported greater confidence as a secondary outcome.

That sentence preserves the hierarchy of evidence.

9. Separate finding strength from explanation strength | Effect 可以强,mechanism 可以弱

A replicated effect does not imply a known mechanism.

You may be able to write:

The performance advantage was replicated across two independent samples.

while still writing:

The process responsible for the advantage remains uncertain.

This combination is mature academic reasoning. Do not let uncertainty about mechanism contaminate confidence in the effect, and do not let confidence in the effect inflate confidence in the mechanism.

10. Interpretation begins by naming the level of claim | 先问“我正在解释什么?”

Possible levels include:

  • existence of an effect;
  • direction;
  • magnitude;
  • duration;
  • mechanism;
  • boundary condition;
  • generalisation;
  • practical significance;
  • theoretical meaning.

One Discussion sentence should not silently jump across all eight levels.

For example, a one-week effect in 96 learners can support a claim about near-term performance. It cannot automatically support a claim about long-term learning, universal applicability, educational policy and cognitive mechanism.

11. Use an interpretation ladder | 建立 interpretation ladder

Start low and climb only when evidence permits:

  1. Observed pattern: Group A scored higher.
  2. Study-level inference: Assignment to A improved the measured outcome.
  3. Construct inference: The intervention may improve independent revision.
  4. Mechanism: The effect may operate through action clarity.
  5. Generalisation: Similar learners may benefit.
  6. Application: Programmes may consider testing this design.

Each step adds assumptions. Each step therefore needs additional evidence or explicit qualification.

12. “This suggests” is not a magic safety phrase | Hedge 不能救坏 inference

Writers sometimes believe adding suggests makes any claim acceptable:

This suggests that the intervention permanently rewires metacognition in all advanced learners.

No. The verb is cautious; the proposition is still enormous.

Calibration requires controlling the whole claim:

The pattern suggests that structured prompts may support near-term revision planning; the study does not establish whether the effect persists or reflects broader metacognitive change.

13. Lesson 014 provides the certainty system | Discussion 是 certainty control 的主战场

Lesson 014 | Calibrate Certainty Without Becoming Vague

Discussion requires the full range:

shows for direct established states;

supports when evidence strengthens a proposition;

suggests for moderate interpretive support;

is consistent with when several explanations remain possible;

may reflect for plausible mechanisms;

cannot rule out for unexcluded alternatives;

remains uncertain when evidence does not resolve the question.

14. “Is consistent with” is especially valuable | 它保留 alternative explanations

If a pattern fits Theory A but also fits Theory B, write:

The result is consistent with Theory A’s prediction.

Do not write:

The result confirms Theory A.

unless the study genuinely distinguishes Theory A from its rivals.

15. A mechanism needs discriminating evidence | Mechanism 不是 plausible story

Suppose structured feedback improves revision and learners report that prompts clarified the next action. It is tempting to write:

The intervention worked because action clarity improved metacognition.

But the study may not have isolated action clarity or measured metacognition.

A stronger Discussion writes:

Learner reports are consistent with the possibility that action clarity contributed to the observed advantage. However, because the intervention also differed in prompting intensity and required revision planning, the present design cannot isolate action clarity as the mechanism.

16. Mechanism claims require more than temporal sequence | “发生在一起”不是机制

A mediator measured after the intervention may correlate with the outcome without mediating the causal effect. Mechanistic language should consider:

  • temporal order;
  • measurement validity;
  • alternative pathways;
  • intervention components;
  • mediator–outcome confounding;
  • experimental manipulation of mechanism where possible.

17. Use rival explanations as a normal part of Discussion | Alternative explanation 不是对自己研究的攻击

Lesson 013 taught counterevidence-resistant argument. The same principle applies here:

Lesson 013 | Build an Academic Argument That Survives Counterevidence

For each main interpretation, ask:

What else could produce this same pattern?

Then compare the explanations instead of pretending yours is the only possible one.

18. Rival explanations can attack different layers | Alternative explanation taxonomy

A result may be explained by:

  • the proposed mechanism;
  • confounding;
  • selection;
  • measurement artefact;
  • regression to the mean;
  • history/time effects;
  • attention/placebo/Hawthorne effects;
  • implementation differences;
  • statistical chance;
  • model specification;
  • attrition;
  • unmeasured subgroup composition.

Not every study faces every rival. The skill is to identify the rivals that genuinely threaten the interpretation you are making.

19. Do not list alternatives mechanically | Rival explanation 也要 prioritise

A Discussion that names ten theoretical possibilities without evaluating them becomes evasive.

Rank alternatives by:

  • plausibility;
  • fit to observed pattern;
  • support from study design;
  • support from prior evidence;
  • ability to explain anomalies;
  • testability.

20. Compare explanations symmetrically | 对自己的解释和 rival 用同一标准

If you criticise a rival because it has no direct mechanism measure, do not accept your preferred explanation without one. If you reject another study because its sample is small, do not ignore the same limitation in your study.

Symmetry is one of the strongest credibility signals in academic writing.

21. Literature comparison is not citation decoration | Discussion 要把 result 放回 scholarly conversation

Elsevier’s guidance recommends situating findings in existing research rather than simply repeating them.

Elsevier | Writing an Excellent Discussion

Useful relationship labels include:

  • replicates;
  • extends;
  • qualifies;
  • contradicts;
  • narrows;
  • provides a boundary condition;
  • offers a possible explanation;
  • fails to reproduce;
  • is consistent with;
  • differs under another population or method.

22. “Consistent with prior research” is too vague alone | Consistent 在哪里?

Weak:

These findings are consistent with previous studies.

Stronger:

The short-term revision advantage replicates earlier reports of improved supported performance, but the weak 12-week difference contrasts with studies that reported persistent effects after repeated retrieval practice.

Now the relationship is specific.

23. Agreement does not prove mechanism | 多个 studies 一致,也可能共享同一 bias

Three observational studies showing the same association strengthen confidence in the pattern but may not resolve causation if all share the same confounding structure.

24. Disagreement is not failure | Contradiction 可能是最有价值的 Discussion material

If your result differs from a major study, ask:

  • population difference?
  • measure difference?
  • intervention dose?
  • follow-up timing?
  • baseline proficiency?
  • implementation fidelity?
  • analysis choice?
  • chance?

25. Explain disagreement without inventing certainty | “可能由于”需要 evidence

Weak:

The difference occurred because our learners were more advanced.

Better:

One possible explanation is baseline proficiency: the present sample was substantially more advanced than the sample in Lee et al., and the intervention may offer less additional support once sentence-level control is already high. This possibility remains speculative because proficiency was not experimentally manipulated.

26. Use chronology carefully | Later study 不自动比 earlier study 更正确

New evidence can:

  • replicate;
  • correct;
  • use better measures;
  • study different contexts;
  • simply disagree.

Do not write a progress story unless the evidence supports one.

27. Explain the result, not your hoped-for result | Discussion 不是 rescue operation

When the primary hypothesis fails, the Discussion should not spend three pages explaining why the intervention “probably still works.”

A disciplined response is:

  1. state the null/uncertain finding;
  2. consider whether the study could detect a meaningful effect;
  3. examine alternative interpretations;
  4. compare with prior evidence;
  5. state what remains unresolved.

28. Null findings require special precision | Null result 不能被当作 “nothing happened”

A null-hypothesis test that is not statistically significant may reflect:

  • no effect;
  • a small effect;
  • low power;
  • high variance;
  • poor measurement;
  • effect heterogeneity;
  • wrong timing.

Discussion should interpret the estimate and uncertainty, not only the binary significance label.

29. Equivalence requires equivalence evidence | “没有差异”不是自动相同

If the study was not designed to test equivalence or non-inferiority, do not write:

The two methods were equally effective.

Write:

The study did not detect a clear between-group difference; the interval remained compatible with effects in either direction.

30. A null primary result cannot be rescued by a positive subgroup without caution | Subgroup rescue 是典型 spin

If a post-hoc subgroup appears positive:

The exploratory subgroup pattern may identify a possible moderator, but it does not overturn the null primary result and requires independent testing.


Part I establishes the Discussion engine: preserve evidence hierarchy, separate finding strength from explanation strength, compare rival explanations fairly, and place findings into the literature without allowing narrative preference to upgrade the data.

Part II — Limitations, bias, generalisability and implication | 第二部分:让限制真正改变 claim

31. Limitations are not a confession | Limitation 不是“研究做得不够好”的道歉

A limitation is a statement about what inference becomes weaker because of the way the study was designed, measured, sampled or analysed. The strongest limitation sentence therefore has two parts:

methodological condition → inferential consequence.

Weak:

The sample was small.

Stronger:

The small subgroup sample produced wide uncertainty around the interaction estimate, so the apparent moderator should be treated as preliminary.

The second sentence tells the reader what must change in the claim.

32. Nature: limitations increase credibility when handled honestly | Limitation 不会自动削弱 paper

Nature Human Behaviour explicitly argues that transparent discussion of limitations does not detract from the value of research and can increase credibility. The danger is not admitting a limitation; the danger is minimising its impact while maintaining the same strong conclusion.

Nature Human Behaviour | Limitations and Appropriate Interpretation

33. Every important limitation should change something | Limitation 必须有 consequence

A limitation can change:

  • certainty;
  • scope;
  • causal interpretation;
  • estimated magnitude;
  • generalisation;
  • mechanism claim;
  • policy recommendation;
  • future research priority.

If the limitation changes none of these, either it is not important or you have not followed its logic far enough.

34. Sample size limitation | 小样本主要影响什么?

Small samples can reduce precision, destabilise subgroup estimates, increase sensitivity to unusual observations and limit the reliability of complex models. But “small sample” does not automatically make all findings invalid.

Write the actual consequence:

The primary estimate remained positive, but the interval was wide enough to include effects from trivial to substantial; the magnitude therefore remains uncertain.

35. Single-site limitation | 单一地点主要影响 transportability

A single-site study can have strong internal validity and narrow external validity.

Weak:

This study was conducted at one school.

Stronger:

Because the trial took place in one high-support programme with experienced instructors, it remains unclear whether similar gains would occur where feedback time and teacher training are more limited.

36. Convenience sample limitation | Convenience sampling 主要影响 representation

If participants volunteered, they may be more motivated, confident or interested than the target population.

Discussion should explain the direction of possible bias if plausible:

Volunteer recruitment may have produced a sample unusually willing to engage with repeated feedback, which could overestimate acceptability in routine settings.

37. Attrition limitation | 谁留下来,会改变 observed effect

When dropout differs between groups or is related to outcome, the final sample can become selective.

Useful Discussion logic:

Attrition was higher among lower-scoring participants in the standard condition. If learners who found the task most difficult were more likely to leave, the observed delayed comparison may underestimate the between-group difference.

Notice that the limitation has a possible direction rather than a generic warning.

38. Missing-data assumptions belong in interpretation | Missing data 处理方法也有 assumption

Multiple imputation, complete-case analysis and model-based methods rely on assumptions. Discussion does not need to reteach the statistical method, but it should recognise when conclusions depend on unverifiable missingness assumptions.

39. Measurement limitation | 你测到了 construct 吗?

If “engagement” is measured only by attendance, the Discussion should not treat the finding as a full measure of cognitive and emotional engagement.

The attendance-based measure captures behavioural participation but not whether learners were cognitively engaged; conclusions should therefore remain limited to observable participation.

40. Proxy outcomes can create construct inflation | Proxy 不要升级成 real-world outcome

Examples:

  • task score ≠ durable learning;
  • self-reported intention ≠ behaviour;
  • click-through rate ≠ understanding;
  • confidence ≠ competence;
  • biomarker ≠ patient-important outcome.

Discussion must preserve the distinction.

41. Self-report limitation | Self-report 不是“低质量 data”,而是回答特定问题

Self-report is appropriate for beliefs, perceived experience, preferences and some subjective states. It is weaker for directly establishing performance or behaviour.

Write:

The survey establishes that participants perceived the feedback as clearer; it does not establish that clarity itself produced the performance difference.

42. Blinding limitation | Knowledge of condition can change behaviour or scoring

When participant or assessor blinding is impossible, discuss the plausible bias pathway rather than simply noting “participants were not blinded.”

43. Implementation fidelity limitation | Intervention 不同于 protocol 本身

A programme may be effective when delivered exactly as intended but difficult to implement routinely. If fidelity varied, Discussion can ask whether outcome heterogeneity followed that variation—but avoid post-hoc mechanism certainty.

44. Short follow-up limitation | Immediate performance ≠ durable capability

This is especially important in learning research.

The one-week advantage demonstrates near-term transfer but does not establish whether the strategy became durable enough to survive a longer period without support.

45. Contamination limitation | Groups can influence one another

If participants share materials or teachers use elements of the intervention with controls, the contrast between conditions narrows. Discussion should consider whether contamination could dilute the estimated effect.

46. Hawthorne and attention effects | Extra attention can be an alternative cause

If one group receives more contact, encouragement or novelty, improved outcomes may reflect attention rather than treatment content. Active comparators help, but if unequal attention remains, say so.

47. Regression to the mean | Extreme baseline selection can create apparent improvement

If participants enter a programme because they performed unusually poorly, later scores may improve partly through ordinary variation. A control/comparison group helps distinguish this effect.

48. History effects | 外部事件可以和 intervention 同时发生

A curriculum change, examination, organisational restructure or software update may affect outcomes during the same period.

49. Maturation | Participants change even without intervention

Children grow, novices practise, patients recover, teams adapt. Time itself can create change.

50. Confounding | Observational Discussion 的核心风险之一

An observed association between X and Y can arise because Z influences both.

Discussion should distinguish:

  • measured confounding addressed by analysis;
  • residual confounding from imperfect measurement;
  • unmeasured confounding that remains possible.

51. “Adjusted for” does not mean “all confounding removed” | Regression adjustment 不是 causal magic

Write:

Adjustment for baseline attainment and study time reduced the association only slightly, but residual confounding by motivation or family support cannot be excluded.

52. Selection bias | Who enters the analysis matters

Selection can occur at recruitment, enrolment, response, follow-up or complete-case analysis. The Discussion should identify which stage could distort the relation being studied.

53. Measurement bias | Measurement error can be directional

If one group knows it received the “new method,” self-reported satisfaction may be inflated. If raters know condition, scoring may shift. If sensors perform differently across groups, measurement itself can create apparent effects.

54. Recall bias | Memory-based reports can differ systematically

Participants with stronger outcomes may remember the intervention more positively. Retrospective questions can therefore entangle outcome and reported exposure.

55. Social-desirability bias | Participants may report what seems expected

This is especially relevant for teacher evaluation, workplace behaviour, health behaviour and sensitive topics.

56. Publication and reporting bias | Your paper sits inside a biased evidence ecosystem

Even if your study is transparently reported, the literature you compare against may overrepresent positive findings. Discussion can acknowledge this when it materially affects claims about “consensus.”

57. Researcher degrees of freedom | Analysis choices can shape findings

Different exclusion thresholds, covariate sets, transformations or subgroup definitions may change results. Sensitivity analyses can show whether conclusions depend on one arbitrary choice.

58. Robustness does not mean truth | Robust to which decisions?

Write:

The primary direction was unchanged across alternative missing-data and outlier specifications.

not:

The finding is robust and therefore true.

59. Pre-registration deviation | Deviations require interpretive honesty

If an outcome or analysis changed after data collection, Discussion should not speak as if the final analysis was always the sole planned test. Distinguish confirmatory and exploratory evidence.

60. Limitations can interact | Two small limitations can combine into a big one

Example:

single site + volunteer recruitment + short follow-up.

Individually each narrows a different dimension. Together they strongly limit claims about routine long-term implementation.

61. Prioritise limitations by impact | 不要列十个 equally

Rank:

  1. threats to primary inference;
  2. threats to magnitude;
  3. threats to scope/generalisation;
  4. minor procedural imperfections.

62. Avoid ceremonial limitation lists | Ritual limitation paragraph

Weak:

This study has several limitations. First, the sample was small. Second, the study was conducted at one site. Third, future research should use larger samples.

Stronger:

The strongest uncertainty concerns persistence rather than the one-week effect: follow-up ended at 12 weeks and attrition reduced the delayed sample. The present data therefore support near-term transfer more strongly than durable capability. In addition, the single high-support site limits transportability to programmes with less intensive feedback infrastructure.

63. Limitations can strengthen the final claim by narrowing it | Narrow = stronger

Original:

Structured feedback produces durable independent writing improvement.

After limitation reasoning:

Structured feedback improved near-term independent revision in this high-support sample; durability and lower-support transportability remain uncertain.

The second claim is more useful because it tells the reader exactly what has been established.

64. Generalisability is not one question | “Can this generalise?” 要拆开

Ask:

  • to which population?
  • to which setting?
  • to which time period?
  • to which implementation conditions?
  • to which outcome?
  • to which version of the intervention?

65. Internal validity and external validity can move independently | 内部强 ≠ 外部广

A tightly controlled trial can strongly estimate an effect in one context while saying little about routine implementation. A large observational dataset may have broad coverage but weaker causal identification.

66. Transportability requires mechanism and context | Context 不只是 location name

To argue transportability, identify which contextual features matter:

  • teacher expertise;
  • class size;
  • baseline proficiency;
  • resource availability;
  • language environment;
  • policy;
  • technology;
  • cultural practice.

67. Do not generalise from “Singapore” or any location as a single mechanism | 地名不是 causal variable

If context matters, specify the actual feature. “This may not generalise outside Singapore” is less informative than “The intervention relied on weekly one-to-one feedback time that may not be available in higher student–teacher ratios.”

68. Population generalisation needs overlap | Who was actually studied?

A trial of advanced adult volunteers does not automatically answer a question about younger learners, beginners or compulsory programmes.

69. Temporal generalisation | Old evidence may travel poorly into changed systems

Technology, curricula, diagnostic criteria and social behaviour change. Discussion should not assume evidence is timeless.

70. Outcome generalisation | One outcome cannot stand in for the entire construct

A revision rubric can estimate revision quality. It does not automatically establish general English proficiency, motivation, creativity or lifelong learning.

71. Mechanism-based generalisation can be stronger than surface similarity | Why should effect travel?

If a mechanism depends on cognitive load, then contexts with similar load conditions may be more relevant than contexts that merely share a country or age label.

72. Practical significance | Statistical evidence must meet real-world scale

Discussion should ask:

  • Is the effect large enough to matter?
  • Compared with normal variation?
  • Compared with cost?
  • Compared with alternative interventions?
  • Does it persist?
  • Does it affect an important outcome?

73. Practical importance requires a benchmark | “Meaningful” needs criterion

Possible benchmarks:

  • minimum clinically important difference;
  • grade boundary;
  • cost threshold;
  • historical change;
  • policy target;
  • expert-defined meaningful difference.

74. Small effects can matter at scale | Effect size needs context

A small individual effect may matter across millions of users if cost is negligible. A moderate effect may be unattractive if implementation is expensive, risky or burdensome.

75. Discussion should distinguish efficacy, effectiveness and implementation | 三个层次不要混

Efficacy: can it work under controlled conditions?

Effectiveness: does it work under routine conditions?

Implementation: can systems deliver it reliably?

A small controlled trial generally cannot answer all three.

76. Implication is not recommendation | Implication ≠ “therefore do it”

Finding:

Near-term transfer improved.

Implication:

Explicit revision actions may be worth testing as a mechanism for transfer.

Recommendation:

All schools should adopt the method.

Each level requires more decision logic.

77. Policy recommendations require criteria beyond effect | Policy adds values and trade-offs

Consider:

  • benefit;
  • cost;
  • risk;
  • equity;
  • feasibility;
  • reversibility;
  • opportunity cost;
  • implementation burden.

78. “Should” is a decision word | 应该 = additional burden of proof

Evidence can support is associated with without supporting should be implemented.

79. Low-risk reversible action can be justified under uncertainty | 不确定不等于不能行动

Example:

Although long-term benefit remains uncertain, the low cost and reversibility of a small pilot make further implementation testing reasonable.

80. High-risk irreversible action requires a higher threshold | Decision threshold depends on consequence

Do not let the Discussion speak as if every implication has the same evidential burden.

81. Theoretical implications need discriminating evidence | Theory A “supported” means what?

A result supports a theory only to the extent that it matches a prediction that relevant rivals do not explain equally well.

82. Do not claim theory confirmation from one compatible result | Compatible ≠ unique

Write:

The pattern is consistent with the retrieval-strength account, but it is also compatible with increased feedback attention; the present design does not distinguish these mechanisms.

83. Theory refinement is often more defensible than theory victory | Boundary conditions are valuable

The effect was limited to novice learners, suggesting that the theory may require baseline automaticity as a boundary condition.

84. Future research should reduce current uncertainty | Future work 不是 “more research is needed”

A strong future-study sentence identifies:

  • the unresolved question;
  • the design needed;
  • the evidence that would discriminate explanations.

85. Turn limitation into a next test | Limitation → future design

Limitation:

No long-term follow-up.

Future test:

A preregistered six-month follow-up with the same unseen-task measure would determine whether the near-term advantage persists after extended unsupported use.

86. Turn rival explanation into a discriminating experiment | Alternative → test

Rival:

Extra feedback attention caused the effect.

Next study:

Match groups on feedback time and word count while varying only action structure.

87. Turn generalisability uncertainty into a transportability test | Context → replication

Replicate across programmes with different class sizes and teacher-feedback capacity.

88. Turn subgroup signal into preregistered moderation | Exploratory subgroup → confirmatory test

Do not simply run more post-hoc subgroup analyses in the same dataset.

89. Future work should not be a wish list | Prioritise the uncertainty that most changes the conclusion

If the biggest uncertainty is duration, do not end by proposing ten unrelated research directions.

90. Replication is a valuable future direction | Novelty is not the only contribution

Independent replication can establish whether a result survives new researchers, contexts and analytic choices.

91. Direct replication and conceptual replication answer different questions | Same procedure vs same theory

Direct replication tests reproducibility of the effect under similar conditions. Conceptual replication tests whether the underlying claim survives a different operationalisation.

92. Multi-site replication tests context sensitivity | One site → network evidence

Variation across sites can reveal whether implementation features or population composition moderate the effect.

93. Discussion should mention what would change your conclusion | Make claims updateable

Advanced academic writing benefits from explicit falsifiability:

A well-powered trial with equal feedback time and no revision advantage would substantially weaken the interpretation that action structure itself drives transfer.

94. Limitations and future work should not contradict each other | Ensure lineage

If limitation says measurement is weak but future work only asks for more participants, you have not solved the core problem.

95. The limitation-to-future-work matrix | 建立 repair map

LimitationWhat it weakensBest next test
short follow-uppersistencedelayed repeated follow-up
single high-support sitetransportabilitymulti-site lower-support replication
unequal contact timemechanism/causationattention-matched comparator
self-report mechanismmechanism validitybehavioural process measure
post-hoc subgroupmoderation claimpreregistered interaction test

96. Bias discussion should be directional when possible | Bias 往哪边推?

Instead of:

Selection bias may be present.

write:

Because higher-motivation learners were more likely to volunteer, the observed adherence and satisfaction levels may be higher than would be expected under compulsory implementation.

97. But do not invent bias direction if unknown | Direction 不清楚就诚实

The direction of any resulting bias is uncertain because dropout differed by both baseline score and study condition.

98. Do not use limitations to dismiss inconvenient results | Limitations apply symmetrically

If you trust the method when the result supports the hypothesis, you cannot suddenly declare the same method invalid when a result is null.

99. Do not minimise limitations with ritual “however” | “Despite these limitations” can become a rhetorical eraser

Weak:

Despite these limitations, the study clearly proves…

Stronger:

These limitations do not affect the direction of the one-week comparison, but they restrict claims about mechanism, durability and routine implementation.

100. The Discussion’s job is to allocate confidence dimension by dimension | Confidence map

Claim dimensionConfidenceReason
one-week directionhighrandomised, precise estimate
magnitudemoderatesingle site
12-week persistencelow–moderatesmaller sample, wide interval
mechanismlownot isolated experimentally
routine implementationlowhigh-support setting

101. This map prevents global hedging | 不要整篇都 “maybe”

Strong where evidence is strong.

Cautious where evidence is weak.

That is C1–C2 propositional precision.


Part II turns limitations into inferential controls. A credible Discussion does not append a weakness list at the end; it lets design, bias, scope and uncertainty actively determine the final claim and the next experiment.

Part III — Discussion architecture across genres and languages | 第三部分:不同 research genre 与中文母语者的高风险点

102. There is no single universal Discussion shape | Discussion 没有一个万能模板

The underlying intellectual jobs recur across disciplines, but the visible structure varies. A biomedical trial, qualitative interview study, computational benchmark paper, systematic review, engineering validation study and humanities article may all “discuss findings” differently.

The C1–C2 skill is therefore functional recognition:

What job is this paragraph doing?

not:

Does this paragraph match the template I memorised?

103. Read target-venue Discussions before drafting | Genre awareness 先于模板

Take three strong recent papers from your target journal or field. Map:

  • where the Discussion starts;
  • whether it opens with a summary;
  • how many findings are discussed;
  • where limitations appear;
  • how literature is compared;
  • whether implications are practical, theoretical or both;
  • whether a separate Conclusion exists;
  • how long the section is relative to Results.

104. Randomised-trial Discussion | RCT Discussion

A strong trial Discussion often prioritises:

  1. primary treatment effect;
  2. secondary outcomes/harms;
  3. comparison with prior trials;
  4. possible mechanisms;
  5. limitations such as adherence, blinding, attrition or generalisability;
  6. clinical/practical implications;
  7. next trial or implementation question.

The trial design may support a strong causal statement about assignment to treatment, but mechanism and real-world implementation may remain uncertain.

105. Observational-study Discussion | Association 不能偷偷变 cause

Observational Discussions must keep causal language under tighter control. Common moves include:

  • describe the association;
  • compare with other observational/experimental evidence;
  • consider confounding;
  • consider reverse causality;
  • consider selection/measurement bias;
  • state what causal interpretation remains plausible rather than established.

106. Cross-sectional Discussion | 时间顺序特别脆弱

If X and Y are measured at the same time, you often cannot know whether X preceded Y.

Write:

The association is compatible with X influencing Y, but reverse causality cannot be excluded because exposure and outcome were measured concurrently.

107. Longitudinal observational Discussion | 时间顺序更清楚,但 confounding 仍可能存在

Repeated observations strengthen temporal reasoning but do not automatically create randomisation.

108. Qualitative Discussion | Meaning, process, context and theory

Qualitative Discussions may integrate interpretation more closely with Findings/Results. They often ask:

  • What process does the theme reveal?
  • How does the meaning compare with prior qualitative work?
  • What contextual feature shapes the pattern?
  • What negative/deviant cases refine the interpretation?
  • How does researcher position affect the account?
  • What conceptual contribution emerges?

109. Qualitative interpretation should remain grounded | “Theme” 不是 imagination licence

Strong qualitative interpretation points back to data patterns, participant accounts, observed interactions or texts. It may be theory-rich, but it should remain traceable.

110. Do not translate qualitative uncertainty into quantitative hedging automatically | 两种 epistemology 不同

“Most participants” or “significant” may be inappropriate if the qualitative design focuses on meaning rather than frequency. Use the conventions of the methodological tradition.

111. Reflexivity belongs where it changes interpretation | Reflexivity 不只是 positionality paragraph

Ask how researcher role, language, institutional position or prior relationships may have shaped data generation and interpretation.

112. Mixed-methods Discussion | Integration is the central job

Mixed-methods Discussion should not simply place a quantitative mini-Discussion next to a qualitative mini-Discussion. It should ask:

  • Where do strands converge?
  • Where do they diverge?
  • Does one explain the other?
  • Does one reveal a boundary the other misses?
  • How does the integrated inference differ from either strand alone?

113. Convergence does not prove mechanism | 两个 strand 一致,也可能共同受 context 影响

If higher scores coincide with reports of clearer action, this supports—but does not prove—the interpretation that action clarity contributed to improvement.

114. Divergence is informative | Quant positive, qual negative 不需要强行调和

Example:

Scores improve, but participants report high cognitive load.

Discussion can interpret this as a benefit–burden trade-off rather than deciding one strand is “right.”

115. Systematic-review Discussion | Synthesis-level reasoning

A systematic-review Discussion often addresses:

  • overall pattern;
  • heterogeneity;
  • risk of bias;
  • certainty/strength of evidence;
  • generalisability;
  • comparison with prior reviews;
  • implications for research/practice.

116. Pooled effect is not the whole story | Meta-analysis 不等于一个 number

If effects vary strongly across populations or designs, Discussion should not present the pooled number as universal.

117. Heterogeneity can be substantive | Variation 可能揭示 boundary conditions

But post-hoc explanations of heterogeneity require caution, especially when based on small numbers of studies.

118. Risk of bias changes conclusion strength | 10 studies ≠ 10 strong studies

If most evidence is high risk of bias, Discussion should reduce certainty even when effect estimates point in one direction.

119. Computational/AI Discussion | Benchmark performance needs external meaning

Questions include:

  • Does performance generalise beyond benchmark?
  • How stable across seeds/models/datasets?
  • What baseline is meaningful?
  • Is improvement practically material?
  • Are there distribution shifts?
  • What failure cases remain?
  • What data leakage or contamination risks exist?

120. Benchmark superiority is not system superiority | 一个 dataset 上赢,不等于 everywhere better

Write scope explicitly.

121. Ablation results can suggest component importance | But “component caused gain” may remain too strong

If removing component X reduces performance, this supports X’s contribution in that model configuration. It does not automatically reveal a universal causal mechanism.

122. Engineering Discussion | Performance, reliability, constraints and deployment

Engineering Discussions often balance:

  • measured performance;
  • operating conditions;
  • failure modes;
  • trade-offs;
  • comparison with existing systems;
  • scalability;
  • safety/reliability.

123. “Better” requires a metric | Faster, cheaper, safer, more accurate, more robust?

Multi-objective systems rarely have one unqualified “best.”

124. Humanities/interpretive Discussion may be distributed across the article | 没有 Discussion heading 也有 discussion work

Interpretive scholarship may integrate evidence and interpretation throughout. The functional equivalents remain:

  • claim;
  • textual/historical evidence;
  • rival reading;
  • context;
  • scope;
  • implication.

125. Do not force empirical vocabulary onto interpretive work | “Results prove” 可能不适合 humanities

Use genre-appropriate verbs:

argues, reads, interprets, traces, demonstrates through textual pattern, situates, complicates.

126. Combined Results and Discussion | 合并 section 合法,但要求更高 signalling

Manchester and many journal conventions allow combined Results/Discussion. When combined, use explicit moves:

Result: The group difference was largest at one week.

Interpretation: This pattern suggests that the benefit may be strongest during early unsupported transfer.

Comparison: The pattern differs from studies reporting stable six-month effects.

Limitation: The present follow-up ended at 12 weeks.

127. Paragraph architecture: F-I-C-R-L-I | 一个实用 Discussion cycle

Use this flexible sequence:

FINDING → INTERPRETATION → COMPARISON → RIVAL → LIMIT → IMPLICATION.

Not every paragraph needs every move, but the model helps diagnose missing reasoning.

128. Finding sentence | 先把重要 result 提升到 conceptual level

The primary finding was a near-term independent-revision advantage concentrated in evidence integration.

129. Interpretation sentence | 回答 “what might this mean?”

This concentration suggests that the intervention affected higher-order revision decisions more strongly than sentence-level correction.

130. Comparison sentence | 放回 prior literature

This pattern extends earlier supported-revision studies by showing a similar advantage on an unseen task after feedback removal.

131. Rival sentence | 提供 alternative

However, because the structured condition also required more explicit planning, additional cognitive engagement could account for part of the difference.

132. Limitation sentence | 限制 inference

The design therefore cannot isolate action structure as the unique mechanism.

133. Implication sentence | 最后才说 next meaning

A component-controlled trial is needed before action structure itself can be treated as the active ingredient.

134. This paragraph is analytical, not verbose | 每句话都推进 inference

Long Discussion does not mean repeating the same cautious idea. Depth comes from testing the interpretation against different constraints.

135. Topic sentences should be claims, not labels | “Feedback” 不是 Discussion topic sentence

Weak:

Feedback is discussed below.

Strong:

The strongest evidence concerns near-term transfer, while durability remains substantially less certain.

136. One paragraph should have one inferential centre | 不要一段同时解释五个 unrelated results

Paragraph coherence improves when each paragraph answers one question:

  • What does finding A mean?
  • Why did finding B differ?
  • How serious is limitation C?

137. Discussion order should reflect intellectual priority | 不一定跟 Results 逐项同顺序

You may begin with the most important finding, then combine related secondary findings, then discuss null results, then limitations.

But every major result should receive the attention its importance warrants.

138. Avoid “citation parade” | Discussion 不是作者名单游行

Weak:

Smith found X. Lee found X. Kumar found Y. Tan found X.

Stronger:

Most prior studies report short-term gains, whereas evidence for persistence is mixed; the present pattern falls between these two states by showing a clear one-week advantage but an imprecise 12-week estimate.

139. Use source clusters by relationship | Citation cluster 应该有逻辑标签

Agreement cluster.

Contradiction cluster.

Mechanism cluster.

Boundary-condition cluster.

140. Do not over-cite obvious interpretation | Citation 应该 support external claims, not every sentence

Your own result does not need a citation to your own paper. Prior literature and external factual claims do.

141. Source recency and foundational status can coexist | 最新 ≠ 唯一重要

Use foundational theory when it defines the debate; use recent work when it represents the current evidence state.

142. Distinguish source claim from your synthesis | 谁说的要清楚

Lee et al. interpret the effect as attentional.

versus:

The combined evidence supports an attentional interpretation.

The second is your synthesis and carries a larger burden.

143. Reporting verbs calibrate source distance | Reporting verb 不只是 style

VerbTypical function
reportsstates a finding
observesdescribes pattern
arguespresents interpretation
proposesoffers model/mechanism
suggestscautious interpretation
demonstratesstrong evidential claim

144. Do not strengthen the source through your verb | Source says “may”; you write “proves”

Preserve the source’s actual certainty unless you have additional justification.

145. Mandarin trap: “说明” | 说明 ≠ always shows/demonstrates

Depending on evidence, English may be:

  • indicates;
  • suggests;
  • is consistent with;
  • shows;
  • supports the interpretation that.

Choose from the evidence relation, not dictionary equivalence.

146. Mandarin trap: “证明” | 证明 ≠ always proves

In informal Chinese academic explanation, 证明 may be used more broadly than English proves. In empirical English writing, proves is often too strong.

Possible reconstructions:

The findings provide evidence that…

The study supports…

The result demonstrates… only when the claim is directly established.

147. Mandarin trap: “由此可见” | From this, it can be seen that… can sound over-certain or unnatural

Better options:

  • Taken together, these findings suggest…
  • The pattern indicates…
  • These results support the interpretation that…
  • The evidence therefore remains consistent with…

148. Mandarin trap: “因此” | therefore encodes logical consequence

Do not use therefore when the relation is merely chronological or suggestive.

Compare:

Scores increased; therefore the intervention caused the improvement.

This is invalid if no causal design supports the step.

149. Mandarin trap: “由于” | because/due to can silently convert explanation into cause

Use:

may reflect

could be related to

one possible explanation is

when mechanism is plausible but not established.

150. Mandarin trap: “可能” | 可能 has several English functions

may = possibility.

might = often weaker/contextual possibility.

is likely to = higher expected probability.

cannot rule out = possibility not excluded, not necessarily likely.

is consistent with = explanation fits evidence but is not unique.

151. Mandarin trap: “应该” | should mixes expectation, recommendation and inference

Discussion must distinguish:

The effect should persist. — expectation.

Researchers should test persistence. — recommendation.

The estimate should be sufficient to… — inference/expectation.

152. Mandarin trap: “显然 / 明显” | clearly/obviously are often rhetorical boosters

Instead of:

Clearly, the intervention was superior.

write the evidence:

The primary estimate favoured the intervention and the interval excluded trivial differences in the prespecified direction.

153. Mandarin trap: “值得注意的是” | notably / it is worth noting that

These phrases can be useful, but often the writer is signalling importance without explaining it.

Replace:

Notably, the delayed effect was smaller.

with:

The delayed effect was smaller, narrowing the evidence for persistence.

154. Mandarin trap: “一定程度上” | to some extent is often vague

Ask what dimension is partial:

  • magnitude?
  • population?
  • duration?
  • mechanism?

Then state the boundary.

155. Mandarin trap: “具有一定意义” | has certain significance is usually too vague

Meaning for theory?

Practice?

Measurement?

Policy?

Name the implication.

156. Mandarin trap: topic-comment omission of source ownership | 谁负责这个 claim?

Chinese can rely on context to carry source ownership. English academic prose often benefits from explicit attribution:

Our data suggest…

Previous trials report…

One interpretation is…

157. Avoid repeated “we believe” | Evidence relationship is more informative

Weak:

We believe the intervention was effective.

Stronger:

The randomised comparison provides evidence of a near-term effect on the measured revision outcome.

158. “We argue” is different from “we prove” | Argument voice

We argue that the pattern is best understood as a boundary effect rather than a complete failure of the mechanism.

This makes clear that you are offering a reasoned interpretation.

159. “We speculate” is deliberately weak | Use for low-evidence possibilities

We speculate that… should not carry the main conclusion of a paper.

160. “We hypothesise” creates a future test | Hypothesis can emerge from Discussion

We therefore hypothesise that the effect depends on baseline automaticity, a prediction that can be tested by stratified randomisation in future work.

161. Discussion language should reveal inferential distance | 距离越远,语言越要显示

Direct:

The group difference was 6 points.

One-step interpretation:

The finding supports a near-term transfer advantage.

Mechanistic interpretation:

The advantage may reflect increased revision planning.

Policy implication:

The result justifies testing the approach under routine conditions.

162. Avoid stacked hedging | “may possibly perhaps” 不是 sophistication

Use one calibrated hedge plus the actual source of uncertainty.

The pattern may reflect action clarity, but the intervention components were not independently manipulated.

163. Avoid stacked boosting | “clearly strongly demonstrates” 也是低 precision

State the design/evidence that earns confidence.

164. Contrast structures are central to Discussion | although / however / whereas / while

These words help preserve multiple truths:

Although the one-week effect was clear, the 12-week estimate was small and imprecise.

Whereas evidence integration improved, sentence-level accuracy changed little.

165. Concession is an intellectual move | Although 不是 style decoration

Concession tells the reader which opposing fact remains true while your broader claim still survives.

166. Use “however” to mark a genuine constraint | 不要每三句一个 however

Too many contrast markers can make prose feel argumentative without improving reasoning.

167. Causal connectors require causal logic | because / therefore / consequently

Use them when the inference is justified, not merely because they create flow.

168. Comparison connectors can express research relationships | similarly / in contrast / unlike / extends

Choose the precise relation rather than generic consistent with.

169. Discussion paragraphs need evidence anchors | 每个 interpretive paragraph 回到哪个 result?

Ask:

Which result is this paragraph interpreting?

If none, the paragraph may be an orphan literature essay.

170. Literature paragraphs need study anchors too | 哪个 finding 需要这段 literature?

Do not write a mini literature review that could have appeared before the Results.

171. The Introduction–Discussion loop | Discussion 应该回到 Introduction 构造的问题

Lesson 017 | Research Introduction

Introduction:

We do not know whether supported gains transfer independently.

Discussion:

The present one-week result provides evidence of near-term independent transfer, while longer-term persistence remains uncertain.

172. If Discussion answers a different question, inspect drift | Paper scope may have changed

Post-hoc findings can be discussed, but they should not replace the primary problem constructed in the Introduction without transparency.

173. HARKing can appear in Discussion too | Result-first theory rewriting

Do not imply that an unexpected result was predicted all along.

174. Protect temporal order of reasoning | planned vs discovered

Use:

Exploratory analyses raise the possibility that…

rather than:

As hypothesised… if it was not hypothesised.

175. Discussion can change the final thesis | Evidence gets the last word

The best Discussion may end with a claim narrower than the Introduction’s hypothesis.

That is not failure. It is evidence-responsive reasoning.

176. The abstract must inherit the final Discussion claim | Abstract certainty cannot stay at pre-analysis optimism

Lesson 016 | Abstract Compression

177. Title must inherit the same claim too | Title 不能继续写 causal claim if Discussion retreats

178. Discussion and Conclusion may be combined or separate | Venue convention matters

If a separate Conclusion exists, Discussion can remain exploratory and detailed while Conclusion compresses.

179. If no separate Conclusion, final Discussion paragraph must perform compression | Final paragraph = calibrated answer

180. Avoid introducing major new literature in the last paragraph | Conclusion should close, not open a new field

181. Avoid last-paragraph policy leaps | Small effect → national policy in one sentence

182. Strong final paragraph has four layers | Final Discussion paragraph

  1. what is established;
  2. what is bounded;
  3. what remains uncertain;
  4. what evidence should come next.

183. Example final paragraph | 示例

Structured action-oriented feedback improved near-term independent revision in this advanced bilingual sample, with the clearest advantage in evidence integration. The study does not establish whether action structure itself was the active mechanism, and the smaller, imprecise 12-week estimate leaves durability unresolved. The single high-support site further limits inference about routine implementation. These findings therefore justify component-controlled and multi-site replication rather than immediate universal adoption.

184. Note how the conclusion gets smaller and stronger | 最终 claim 不是最大,而是最稳

It does not claim:

  • universal effectiveness;
  • known mechanism;
  • durability;
  • policy mandate.

It does claim what the evidence earns.


Part III turns Discussion from a generic “interpretation section” into genre-aware advanced English. It also separates Mandarin rhetorical habits from English evidential commitments so that translation does not silently inflate causality, certainty or recommendation strength.

Part IV — Full worked Discussion, diagnostics and independent transfer | 第四部分:完整案例、诊断与独立迁移

185. The complete fictional study | 完整虚构研究状态

All data and studies in this worked case are fictional teaching material.

Research question:

Does structured action-oriented feedback improve independent revision among advanced bilingual learners after feedback is removed?

Design:

96 advanced bilingual learners were randomly assigned 1:1 to structured action-oriented feedback or standard evaluative feedback. Both groups received equal feedback time and similar word counts. The primary outcome was an unseen independent revision task one week later. A secondary follow-up occurred at 12 weeks. Two blinded raters scored claim clarity, evidence integration, reasoning, organisation and sentence accuracy.

Primary result:

The structured group scored 6.0 points higher at one week (95% CI 2.4–9.6).

Rubric pattern:

The largest difference occurred in evidence integration; reasoning also improved; sentence-level accuracy differed little.

12-week result:

The structured group remained 2.1 points higher, but the estimate was imprecise (95% CI −1.2 to 5.4).

Qualitative strand:

Participants in the structured condition frequently described the prompts as making the “next revision action” clearer. A minority reported that the prompts felt restrictive.

Context:

Single high-support programme with experienced instructors.

186. Weak Discussion version | 弱版本

The results clearly prove that structured feedback is superior to ordinary feedback. The intervention helped students think more deeply because the action prompts improved metacognition. This finding is consistent with previous research and demonstrates that learners need clear structure. Although the study was small and conducted in one programme, the results are highly significant and have important implications. Schools should therefore adopt structured feedback widely. Future research should use larger samples.

187. Diagnose every failure in the weak version | 逐项诊断

“clearly prove” overstates certainty.

“superior” is unbounded: superior on which outcome, time horizon and cost?

“think more deeply” is not the measured construct.

“because” converts a plausible mechanism into an established cause.

“improved metacognition” introduces an unmeasured construct.

“consistent with previous research” gives no specific relationship.

“learners need clear structure” generalises beyond the sample and intervention.

“highly significant” confuses statistical and practical meaning.

“schools should therefore adopt” makes a policy leap.

“larger samples” does not target the largest uncertainties: mechanism, durability and transportability.

188. Step 1: Write the primary-answer sentence | 先回答研究问题

Structured action-oriented feedback improved near-term independent revision in this sample, with the clearest advantage in evidence integration.

Why this works:

  • “improved” is supported by random assignment to condition;
  • “near-term” preserves the one-week horizon;
  • “in this sample” controls scope;
  • “evidence integration” identifies where the effect concentrated.

189. Step 2: Interpret the pattern without inventing the mechanism | Pattern → cautious meaning

The concentration of the effect in evidence integration rather than sentence accuracy suggests that the intervention influenced higher-order revision decisions more strongly than local language correction.

This sentence remains close to measured rubric dimensions. It does not yet claim metacognition, cognitive load, motivation or a specific neural/cognitive mechanism.

190. Step 3: Bring in the qualitative evidence carefully | Qualitative strand adds plausibility

Participant accounts are consistent with this interpretation: learners frequently described the action prompts as clarifying what to change next. However, these reports identify perceived usefulness rather than a directly measured cognitive mechanism.

Now the qualitative strand supports an interpretation while preserving its evidential role.

191. Step 4: Compare with prior literature by relationship | Literature relation

Imagine earlier fictional studies:

  • Study A found better supported revision with action prompts.
  • Study B found no delayed benefit after six months.
  • Study C found stronger effects in novice learners.

A strong sentence:

The one-week transfer advantage extends earlier evidence from supported revision tasks, whereas the smaller 12-week estimate is more consistent with reports that feedback-related gains weaken after structured support is removed for longer periods.

192. Step 5: Expose the strongest rival explanation | Rival before celebration

Even with equal feedback time and word count, the structured condition required more explicit planning. Therefore:

The present design does not isolate whether the benefit resulted from action structure itself or from the additional planning demands imposed by the prompts.

This rival is stronger than a generic objection such as “students might have tried harder.” It points to a concrete component difference.

193. Step 6: Use the delayed result to narrow the claim | 12-week evidence changes thesis

The 12-week estimate was smaller and imprecise, so the study supports near-term transfer more strongly than durable independent capability.

Notice that the later evidence is not hidden. It actively changes the conclusion.

194. Step 7: State the external-validity limitation | Single-site consequence

Because the trial was conducted in a high-support programme with experienced instructors, the findings may not transport directly to settings where feedback time and instructor preparation are more limited.

195. Step 8: State the strongest practical implication | Implication, not mandate

The findings justify further testing of action-oriented feedback under controlled component comparisons and more routine implementation conditions; they do not yet justify universal replacement of standard feedback.

196. Complete model Discussion paragraph 1 | Primary finding cycle

Structured action-oriented feedback improved near-term independent revision in this advanced bilingual sample, with the clearest advantage in evidence integration. The concentration of the difference in higher-order rubric dimensions rather than sentence accuracy suggests that the intervention affected revision decisions more strongly than local correction. Participant accounts are consistent with this interpretation, as learners frequently described the prompts as clarifying the next action. However, these reports measure perceived usefulness rather than the cognitive mechanism itself, and the structured condition also required more explicit planning than the comparison condition. The design therefore cannot isolate action structure as the unique active component.

197. Complete model Discussion paragraph 2 | Literature + persistence cycle

The one-week transfer advantage extends earlier findings from supported revision tasks by showing that a performance difference remained when learners revised an unseen text without direct feedback. Evidence for persistence was weaker. At 12 weeks the estimated group difference was smaller and imprecise, a pattern that is compatible with previous reports of declining feedback-related gains after support is removed for longer periods. The present study therefore distinguishes near-term transfer from durable capability: it provides relatively strong evidence for the former but leaves the latter unresolved.

198. Complete model Discussion paragraph 3 | Boundary + implication cycle

Generalisation is further limited by the study setting. The trial was conducted in one high-support programme in which instructors were experienced with the feedback protocol and learners were already highly proficient. These conditions may not represent settings with larger groups, lower feedback capacity or less advanced learners. The findings should therefore be treated as evidence that structured action-oriented feedback can support near-term transfer under favourable implementation conditions, not as evidence that the approach will produce the same effect across educational systems. Multi-site replication with variation in feedback capacity would provide a stronger test of transportability.

199. Complete model final paragraph | Calibrated close

Taken together, the findings support a near-term independent-revision benefit from the structured feedback condition, concentrated in evidence integration and reasoning. The study does not establish the specific mechanism, long-term durability or routine-setting effectiveness of the approach. A component-controlled replication that matches planning demands, followed by longer multi-site testing, would determine whether action structure itself produces a durable and transportable benefit. Until then, the evidence supports further targeted testing rather than universal adoption.

200. Why the full model works | 它没有把研究“说小”,而是说准

The final claim is narrower than the weak version but more useful. It tells readers:

  • what is established;
  • where the effect appears;
  • what mechanism remains unresolved;
  • what duration remains unresolved;
  • what context limits generalisation;
  • what next design would reduce uncertainty.

201. The Discussion matrix | 写之前先做矩阵

FindingInterpretationRivalPrior literatureLimitNext test
6-point one-week gainnear-term transferplanning demandextends supported-task studiesmechanism not isolatedcomponent control
evidence integration strongesthigher-order effectrubric sensitivitymatches action-oriented feedback workconstruct boundaryindependent measure
12-week estimate weakdurability uncertainattrition/noisematches fading-effect studiesprecisionlarger delayed follow-up

202. This matrix prevents “story first” writing | 先让每个 claim 有 evidence parent

When the matrix is built before prose, it becomes harder to hide contradictory results or invent mechanisms midway through a paragraph.

203. The primary-finding audit | 主 finding audit

Answer:

  1. What was the primary question?
  2. What result directly answers it?
  3. How precise is the estimate?
  4. What is the strongest defensible conceptual restatement?
  5. What does the result not establish?

204. The explanation audit | Explanation audit

For each “because”, “due to”, “reflects”, “explains”, or “mechanism” sentence, ask:

  1. Was this variable/process measured?
  2. Did it precede the outcome?
  3. Was it manipulated or merely observed?
  4. What rival mechanism fits the same pattern?
  5. What evidence would distinguish them?

205. The prior-literature audit | Literature relation audit

For every citation cluster, label:

replicate / extend / qualify / contradict / boundary / mechanism / context.

If you cannot label the relationship, the citation may be decorative.

206. The null-result audit | Null honesty audit

Ask:

  • Which result most challenges my preferred narrative?
  • Where is it discussed?
  • Does it change the conclusion?
  • Did I falsely convert non-significance into no effect?

207. The limitation-consequence audit | Every limitation earns a consequence

LimitationConsequence stated?
small sampleprecision?
single sitetransportability?
self-reportconstruct/mechanism?
short follow-uppersistence?
unblinded ratingmeasurement bias?

208. The “despite” audit | 是否用一个 despite 把 limitation 全擦掉?

Search for:

despite these limitations

Then ask whether the conclusion immediately returns to its original maximum strength. If yes, the limitation paragraph may be ceremonial.

209. The generalisation audit | Scope map

Highlight every population/context word in the conclusion:

learners, students, patients, users, schools, companies, systems.

Replace broad nouns with the actual studied boundary where necessary.

210. The mechanism audit | Mechanism map

List every causal process the Discussion names. For each, classify:

  • directly measured;
  • indirectly supported;
  • plausible only;
  • speculative.

211. The recommendation audit | Should / must / ought

Every recommendation should have:

  • evidence;
  • goal;
  • cost/risk consideration;
  • scope;
  • decision threshold.

212. The “what would change my mind?” audit | 可更新性

Write one sentence:

My conclusion would weaken if ______.

If nothing could weaken it, the Discussion may be protecting a belief rather than interpreting evidence.

213. The alternative-explanation audit | 至少写两个 rivals

For each major mechanistic interpretation, list:

  1. preferred explanation;
  2. strongest rival;
  3. second plausible rival;
  4. evidence currently favouring one;
  5. future test that distinguishes them.

214. The symmetry audit | Same standards for preferred and rival explanation

Would you accept the same quality of evidence if it supported the opposite conclusion?

215. The abstract-consistency audit | Abstract must match final Discussion

Compare:

  • main claim;
  • causal verb;
  • time horizon;
  • population;
  • limitation;
  • recommendation.

216. The title-consistency audit | Title can become spin

If Discussion concludes “associated with”, title cannot say “causes”.

If Discussion says “one-week improvement”, title cannot say “lasting learning”.

217. The citation-ownership audit | Source voice vs writer voice

Mark:

source says

our data show

we interpret

field consensus suggests

Do not blend them.

218. The Mandarin back-translation audit | 中文回译可以暴露 certainty drift

Take your English sentence:

The findings are consistent with the possibility that action clarity contributed to the effect.

Translate it back into Chinese.

If your internal version becomes:

结果证明作用机制就是 action clarity。

you have changed the claim.

219. The Chinese→English reconstruction drill | 从 evidence state 重建,不从字面翻译

Chinese note:

结果说明结构化反馈在一周后仍然有效,但十二周后差异不明显,因此可能是短期作用。

Literal-risk version:

The results prove that structured feedback was still effective after one week, but there was no difference after twelve weeks; therefore the effect is short term.

Calibrated reconstruction:

The one-week data support a near-term benefit from structured feedback. At 12 weeks the estimated difference was smaller and imprecise, so the present study does not establish persistence beyond the shorter follow-up period.

220. Why the reconstruction is better | 三个主要修复

证明 → support.

没有差异 → smaller and imprecise estimate.

因此是短期作用 → does not establish persistence.

The English version preserves uncertainty rather than forcing a binary story.

221. Drill: “这说明” | Translate by evidence relation

Case A: direct table value.

This shows that… can be appropriate.

Case B: causal interpretation from observational data.

This suggests that… or is consistent with… may be more accurate.

222. Drill: “可能由于” | Distinguish plausible cause from not-excluded cause

may reflect = plausible explanation.

cannot rule out = possibility remains, but not necessarily supported.

These are not interchangeable.

223. Drill: “应该” | Separate expected, recommended, required

The effect should be replicated can mean recommendation.

The effect should persist can mean expectation.

Participants should meet the criteria can mean requirement.

224. Drill: “显著” | Statistical vs substantial

Chinese “显著” may mean noticeable, marked or statistically significant depending context.

English Discussion must choose:

  • statistically significant;
  • substantial;
  • large;
  • clear;
  • marked;

Do not let one word carry incompatible meanings.

225. Practice A — Primary result | 练习 A

Result: intervention group 5 points higher; randomised design; one-week task.

Write the first Discussion sentence.

226. Practice B — Mechanism | 练习 B

Interviewees say prompts helped them plan. Planning was not measured behaviourally.

Write one cautious mechanism sentence.

227. Practice C — Rival explanation | 练习 C

Intervention required two extra minutes of planning.

Write the strongest rival explanation.

228. Practice D — Prior literature | 练习 D

Earlier studies show immediate improvement but mixed delayed effects.

Relate your one-week positive / twelve-week imprecise pattern.

229. Practice E — Null result | 练习 E

12-week estimate 1.1, 95% CI −2.2 to 4.4.

Write a Discussion interpretation without claiming equivalence.

230. Practice F — Limitation consequence | 练习 F

Only one high-resource site.

Write what inference is weakened.

231. Practice G — Generalisation | 练习 G

Participants are advanced adult volunteers.

Repair:

Structured feedback helps language learners.

232. Practice H — Policy implication | 练习 H

Small benefit; low-cost intervention; no harm signal; only one trial.

Write a proportionate recommendation.

233. Practice I — Subgroup | 练习 I

Post-hoc effect larger among low-baseline learners.

Write Discussion language that preserves exploratory status.

234. Practice J — Limitation future test | 练习 J

No mechanism measure.

Write a future-study design that would reduce uncertainty.

235. Practice K — Qualitative divergence | 练习 K

Most participants report clarity; minority report restriction.

Write a balanced interpretation.

236. Practice L — Observational causality | 练习 L

Tool use associated with higher grades; motivation not measured.

Write a causal-boundary sentence.

237. Practice M — Theory | 练习 M

Finding matches Theory A and Theory B.

Repair:

The study confirms Theory A.

238. Practice N — “由此可见” | 练习 N

Reconstruct without using it can therefore be seen that.

239. Practice O — Conclusion | 练习 O

Write four sentences:

  1. established;
  2. bounded;
  3. uncertain;
  4. next evidence.

240. Model answers | 示范答案

A: Assignment to structured feedback improved one-week independent revision in this sample.

B: Participant accounts are consistent with the possibility that clearer planning contributed to the benefit, although planning was not measured directly.

C: The additional planning time may account for part of the observed advantage, so the present design cannot isolate feedback structure as the sole active component.

D: The one-week advantage extends previous evidence for immediate gains, while the smaller 12-week estimate is consistent with the less stable delayed effects reported elsewhere.

E: The 12-week data do not establish a persistent benefit; the interval remains compatible with effects ranging from a small disadvantage to a modest advantage.

F: The high-resource setting limits transportability to programmes with less feedback time or instructor preparation.

G: The findings support a near-term benefit among the advanced adult volunteers studied; effects in younger, beginner or compulsory-learning populations remain uncertain.

H: The low cost and reversibility of the approach justify further pilot testing, but one trial is insufficient to support routine universal adoption.

I: An exploratory post-hoc pattern suggested a larger effect among lower-baseline learners; this possible moderation requires preregistered replication.

J: A component-controlled trial could measure planning behaviour directly while holding feedback time constant, allowing the proposed mechanism to be tested.

K: Action structure appeared to improve clarity for many participants, but reports of restriction indicate that the same structure may become constraining once learners already possess stable revision plans.

L: Higher tool use was associated with higher grades, but unmeasured motivation could influence both variables and prevents a firm causal interpretation.

M: The finding is compatible with both Theory A and Theory B and therefore does not distinguish between them.

241. The 20-minute Discussion drill | 20 分钟训练

TimeTask
3 minwrite primary-answer sentence
4 minlist two interpretations + two rivals
4 minconnect finding to prior literature
4 minwrite limitation → consequence
5 minwrite calibrated final paragraph

242. The 45-minute growth session | 45 分钟增长模式

TimeTask
7 minevidence-confidence map
8 minDiscussion matrix
10 mintwo finding cycles
8 minlimitations + future test
7 minMandarin-English certainty audit
5 minfinal claim compression

243. The 90-minute deep session | 90 分钟深度模式

TimeTask
10 minreconstruct Results hierarchy
15 minmap literature relationships
15 minrank rival explanations
15 minwrite three Discussion cycles
10 minlimitation/bias consequences
10 minimplication/future study
15 minfull audits + final paragraph

244. Seven-day Discussion cycle | 七天训练循环

  1. Day 1: result restatement vs repetition.
  2. Day 2: interpretation levels and certainty.
  3. Day 3: prior literature relationships.
  4. Day 4: rival explanations and mechanism tests.
  5. Day 5: limitations, bias and generalisation.
  6. Day 6: implication, future work and final claim.
  7. Day 7: unseen full Discussion benchmark.

245. Twelve-week C1–C2 Discussion progression | 12 周路线

WeeksFocusOutput
1–2finding → interpretationcalibrated finding paragraphs
3–4prior literature + relationship languagecomparison cycles
5–6rival explanation + mechanismmodel-comparison paragraphs
7–8limitations + bias + scopeconsequence-bearing limitation sections
9–10implications + future research + conclusionscomplete Discussions
11–12genre and bilingual transfersubmission-ready portfolio

246. Monthly benchmark | 每月基准任务

Choose one research paper whose Results you can access and understand. Without reading its Discussion first:

  1. write the primary finding;
  2. rank confidence across direction, magnitude, mechanism, scope and persistence;
  3. propose two interpretations;
  4. write two rival explanations;
  5. identify the prior literature relationship;
  6. write three limitations with consequences;
  7. write one generalisation boundary;
  8. write one practical implication;
  9. design one discriminating future study;
  10. write a 600–900 word Discussion;
  11. then read the published Discussion;
  12. compare what the authors inferred with what you inferred;
  13. identify where either version overclaims;
  14. reconstruct the final claim in Simplified Chinese;
  15. rebuild it in English and check certainty drift.

247. The one-page Discussion operating system | 一页 Discussion 操作系统

  1. Answer the primary question first.
  2. Restate conceptually; do not repeat numbers.
  3. Name the level of inference. Effect / magnitude / duration / mechanism / scope.
  4. Calibrate certainty dimension by dimension.
  5. Connect to prior research by relationship.
  6. Steelman the strongest rival explanation.
  7. Do not rescue null findings with secondary positives.
  8. Turn every important limitation into an inference consequence.
  9. Separate internal validity from generalisation.
  10. Separate implication from recommendation.
  11. Match recommendation strength to risk/cost/reversibility.
  12. Turn the biggest uncertainty into a discriminating next study.
  13. Keep source voice, data voice and writer interpretation distinct.
  14. Audit Mandarin→English certainty and causality drift.
  15. Make the abstract/title inherit the final calibrated claim.
  16. End with the strongest defensible claim, not the largest possible claim.

248. Research and reference floor | 研究与参考基础

249. Canonical eduKate routes | eduKate 主页面路由

250. SEO language map | 本课自然覆盖的搜索意图

This lesson naturally serves learners searching for how to write discussion section, research discussion section, discussion vs results, interpret research findings, limitations section, research limitations, alternative explanations, generalisability, implications of findings, future research, cautious academic language, overclaiming research, how to discuss null results, how to compare findings with literature, C1 academic writing, C2 academic English, academic English for Chinese speakers, English for Mandarin speakers, Discussion 怎么写, 研究讨论怎么写, 研究限制怎么写, 结果解释, 研究意义, future research 怎么写, 学术英语 Discussion, 中文母语学术英语 and C1 C2 research writing.

251. Final assignment | 最终作业

Choose a research paper or a permitted/fictional dataset with at least one main finding, one secondary or null finding and enough methodological detail to evaluate limitations.

Complete this chain:

  1. Write the primary research question.
  2. Write the primary finding in one sentence without interpretation.
  3. Write the strongest defensible interpretive restatement.
  4. Identify the exact claim level: direction, magnitude, duration, mechanism, boundary or generalisation.
  5. Write two plausible explanations.
  6. Write the strongest rival explanation.
  7. State which evidence currently favours each explanation.
  8. Describe how the finding relates to at least three prior studies by relationship, not author order.
  9. Discuss the most important null or contradictory result.
  10. Write three limitations, each with an inference consequence.
  11. Rank confidence separately for effect, magnitude, persistence, mechanism and generalisation.
  12. Write one practical implication without using should.
  13. Write one recommendation, then justify the decision threshold, cost/risk and scope.
  14. Design one future study that discriminates between the preferred and rival explanation.
  15. Write a full Discussion of 1,500–2,500 words.
  16. Run the null-honesty audit.
  17. Run the limitation-consequence audit.
  18. Run the symmetry audit.
  19. Run the Mandarin back-translation audit.
  20. Rewrite the abstract and title so they inherit the final calibrated claim.

Then ask the hardest question:

If the most inconvenient result in my study appeared in the first paragraph of the Discussion, would the rest of my reasoning still survive?

如果把这项研究里最不方便、最不支持我原来故事的 finding 放在 Discussion 第一段,我后面的 argument 还能不能站得住?

If yes, the Discussion is probably interpreting evidence rather than protecting a preferred narrative.

If no, revise the narrative.


Continue | 继续

The next lesson will focus on the final compression problem of the full research paper: how to write a conclusion that closes the argument without merely repeating the abstract, inventing certainty, or opening a new claim at the last moment.

下一课进入 research paper 最后的 compression gate:怎样写 Conclusion,让整篇 argument 真正关闭,同时不重复 abstract、不突然提高 certainty,也不在最后一段打开新的 claim。

Next: EDKS-ADV-ZH-0021 · Lesson No.021 · Write a Conclusion That Closes the Argument Without Making It Bigger | 把研究结论写到真正收束,而不是最后再扩大主张

Back to Advanced English Chinese Edition Hub · 返回高级英语中文版主页

Part V — Peer-review survival, spin control and advanced transformation laboratory | 第五部分:让 Discussion 经得住 reviewer 与最终发表

252. A strong Discussion must survive a hostile but fair reader | Reviewer test

The final stage of advanced Discussion writing is not asking whether the prose sounds persuasive. It is asking whether a knowledgeable reader who actively looks for overreach can still accept the reasoning structure.

A fair reviewer may ask:

  • Did the authors answer the question they actually designed the study to answer?
  • Did they distinguish primary from exploratory findings?
  • Did they turn association into causation?
  • Did they acknowledge the strongest alternative explanation?
  • Do limitations change the final claim?
  • Is generalisation wider than the sample and setting permit?
  • Does the abstract/title make a stronger claim than the Discussion?

These are not attacks on the paper. They are tests of whether the article has preserved the chain from evidence to inference.

253. Reviewer criticism often identifies a broken inference, not a bad sentence | Reviewer 不是 grammar checker

A reviewer comment such as “the authors overstate the implications” rarely means the adjective choice alone is wrong. It may mean:

  • the study measured a proxy but the conclusion names the target construct;
  • the study was observational but the Discussion uses causal verbs;
  • the sample is narrow but the conclusion is universal;
  • the main result is uncertain but the conclusion is categorical;
  • a positive secondary outcome has displaced the primary outcome.

Repair the inference first. Then repair the language.

254. Build a reviewer-response map before rewriting | Comment → underlying problem → repair

Reviewer commentLikely problemBest repair
“Overstated causal language”design–claim mismatchrestore association language or add causal justification
“Mechanism not established”plausible story presented as causeseparate mechanism hypothesis from effect
“Generalisation too broad”population/context inflationrestore sample/setting boundary
“Limitations superficial”no inferential consequencestate what each limitation weakens
“Discussion too repetitive”Results restated instead of interpretedcompress numbers; expand comparison/rival/limit

255. Do not defend every original sentence | Revision is not litigation

The purpose of a response letter is not to prove the first draft was secretly correct. If a reviewer identifies a real inferential problem, change the paper.

Weak response mindset:

We respectfully disagree because our interpretation is reasonable.

Better:

We agree that the original wording exceeded the design. We have replaced causal language with association language and added the residual-confounding limitation in the Discussion.

256. Distinguish “reviewer preference” from “evidential correction” | 两类 revision 不一样

Some comments are stylistic or journal-specific:

  • move limitations earlier;
  • shorten Discussion;
  • combine two paragraphs;
  • use a separate Conclusion.

Other comments concern truth conditions:

  • causal overclaim;
  • misstated null result;
  • hidden attrition;
  • unsupported subgroup;
  • misrepresented prior study.

Fix truth conditions first.

257. Spin is often created by hierarchy, not false numbers | Spin 不一定是“造假”

A paper can report every number correctly and still mislead by deciding which number receives narrative emphasis.

Common forms:

  • primary null result buried; secondary positive result foregrounded;
  • short-term benefit headlined; long-term null omitted;
  • relative effect emphasised; absolute difference hidden;
  • subgroup result highlighted without interaction evidence;
  • harms placed in supplement while benefits appear in abstract;
  • observational association described with intervention language.

258. The spin audit begins with hierarchy | What was the study’s main test?

Write, before reading the Discussion:

Primary question → primary outcome → primary estimate → primary uncertainty.

Then compare that four-part state with:

  • Discussion first paragraph;
  • final paragraph;
  • abstract;
  • title;
  • press-release style summary if one exists.

259. Positive framing can be mathematically true and rhetorically misleading | framing effect

Example:

The intervention doubled success.

If success rose from 1% to 2%, the relative statement is correct but incomplete for practical interpretation.

A Discussion may need:

Success increased from 1% to 2%, corresponding to a doubling in relative terms but an absolute increase of one percentage point.

260. Negative framing can mislead too | Same evidence, opposite story

Half of participants failed to improve.

versus:

Half of participants improved.

Use the framing that answers the research question, and provide the underlying proportion so readers can reconstruct the state.

261. Headline verbs deserve a separate audit | Title/abstract verbs are high leverage

Audit:

  • causes;
  • prevents;
  • improves;
  • predicts;
  • is linked to;
  • is associated with;
  • may support;
  • fails to;

The same data can be made much stronger or weaker by verb choice. The verb must inherit the design.

262. “Predicts” has two meanings | Statistical prediction vs temporal prediction

In modelling, a variable may predict an outcome statistically without being causal. In ordinary English, readers may hear “predicts” as forecasting the future.

Discussion should clarify the intended meaning where ambiguity matters.

263. “Leads to” is usually causal | Do not use as a synonym for “comes before”

Higher motivation led to higher grades is a causal claim.

If the study observed an association:

Higher motivation was associated with higher grades.

264. “Explains” is stronger than “accounts for variance” | Statistical explanation vs causal explanation

A regression model that accounts for 30% of variance does not prove that the included variables are the causal explanation.

265. “Mediates” is a technical causal claim in many fields | Mediation needs design/assumptions

Do not use mediates simply because M correlates with X and Y.

266. “Moderates” requires an interaction relationship | Subgroup difference vs moderator

Different significance levels across subgroups do not establish moderation. The interaction itself must be evaluated.

267. “Independent predictor” does not mean independent cause | Adjustment is not randomisation

In observational models, “independent predictor” often means statistically associated after included covariates—not causally independent of all other pathways.

268. “Risk factor” can be descriptive or causal depending field | Clarify if necessary

If causality is not established, risk marker or associated factor may sometimes be more precise.

269. The counterfactual test for causal language | What would have happened otherwise?

Before using a strong causal phrase, ask:

Does the study design support a comparison between what happened and what would have happened to the same target under the alternative condition?

Randomisation, natural experiments, strong quasi-experimental designs and causal models address this in different ways; simple association usually does not.

270. The construct test for broad nouns | What exactly was measured?

Highlight nouns such as:

  • learning;
  • intelligence;
  • wellbeing;
  • engagement;
  • performance;
  • quality;
  • success.

Then replace them with the actual operational outcome where necessary.

271. The time-horizon test | When is the claim true?

Improves retention means little without a time scale.

Immediate?

One week?

Six months?

Five years?

272. The comparator test | Better than what?

The programme was effective.

Compared with:

  • no treatment?
  • business as usual?
  • another active intervention?
  • baseline?
  • a historical control?

Discussion must preserve the comparator because “effective” depends on it.

273. The dose test | Which version of the intervention?

Twenty minutes once is not the same intervention as one hour weekly for twelve weeks. Generalisation should preserve dose where effect may depend on it.

274. The implementation test | Who delivered it and with what skill?

A complex intervention delivered by its designers may perform differently when delivered routinely by less specialised staff.

275. The failure-mode test | What happened when the intervention did not work?

Strong Discussions examine:

  • non-responders;
  • adverse cases;
  • implementation failures;
  • boundary conditions;
  • unexpected direction changes.

Failure patterns often reveal more about mechanism than average success.

276. Average effects can conceal opposing responses | Heterogeneity matters

An average improvement of 3 points can arise from:

  • everyone improving 3 points;
  • half improving 10, half declining 4;
  • novices improving strongly, advanced learners not changing.

Discussion should not treat these as equivalent if subgroup/individual patterns are relevant and appropriately analysed.

277. Do not invent subgroups after the fact to rescue heterogeneity | subgroup mining

Exploratory heterogeneity can generate hypotheses. It should not be presented as established moderation without replication.

278. Unexpected direction deserves a mechanism audit | Opposite result

If an intervention expected to help appears harmful:

  1. verify coding/analysis;
  2. inspect measurement;
  3. consider implementation;
  4. consider plausible harm mechanism;
  5. compare prior evidence;
  6. do not hide the result.

279. Ceiling and floor effects can distort interpretation | Measurement range matters

Advanced participants may show little gain because the instrument cannot register improvement near its maximum. Very weak participants may show little decline if scores cluster at the floor.

280. Practice effects can mimic learning | Repeated tests may improve scores

If participants take similar items repeatedly, Discussion should consider whether familiarity rather than underlying capability changed.

281. Demand characteristics can influence behaviour | Participants infer the study aim

This may matter especially in behavioural, educational and social research where participants know the intervention’s intended direction.

282. Instructor expectancy can influence outcomes | Researcher/teacher expectations

If teachers know which group receives the “new” intervention, encouragement or scoring behaviour may differ.

283. Contamination can shrink contrasts | Control group learns the intervention

When participants interact, control members may adopt treatment strategies, reducing apparent effect size.

284. Non-adherence changes estimand interpretation | Assignment vs treatment received

Intention-to-treat estimates the effect of assignment under real adherence. Per-protocol analyses answer a different question and may reintroduce selection.

285. Mechanism from process data | Process measures can strengthen, not guarantee, mechanism

If planning behaviour increases before performance improves, and mediation assumptions are plausible, the mechanism case strengthens. But causal mediation still requires strong assumptions.

286. Triangulation strengthens inference when errors differ | Multiple methods

Survey, behavioural measure and observational record agreeing can increase confidence if they do not share the same bias source.

287. Triangulation does not mean “three methods = truth” | Shared bias can remain

Three self-report instruments administered in the same context may reproduce the same response bias.

288. External evidence can strengthen the Discussion | Converging evidence beyond one study

A small trial’s mechanistic interpretation may become more plausible when:

  • laboratory studies support the process;
  • longitudinal studies show the expected temporal relation;
  • qualitative work documents the experience;
  • replications reproduce the effect.

But the Discussion should still distinguish what this study establishes from what the broader literature supports.

289. The Discussion is a model-selection exercise | Which story explains the most with the least distortion?

Compare candidate explanations by:

  • fit to primary result;
  • fit to null/negative results;
  • fit to subgroup variation;
  • compatibility with prior evidence;
  • assumption burden;
  • predictive discriminability.

290. Prefer explanations that survive inconvenient evidence | Anomaly test

If your preferred mechanism explains the positive one-week result but cannot explain the disappearing 12-week effect, you need either:

  • a time-dependent mechanism;
  • a boundary condition;
  • or a different explanation.

291. A Discussion can legitimately end with unresolved alternatives | Not every paper needs closure

The present evidence does not distinguish between increased planning effort and action clarity as mechanisms.

This is a useful scientific conclusion when true.

292. Unresolved does not mean uninformative | Partial knowledge is still knowledge

You may know:

  • that an effect exists;
  • that it is near-term;
  • that it concentrates in one outcome;
  • that two mechanisms remain plausible.

That is a meaningful state.

293. Mini-case 1 — Education | Positive immediate, null delayed

Finding:

Retrieval intervention improves same-week quiz but not three-month test.

Weak Discussion:

Retrieval practice is effective.

Calibrated:

The intervention improved short-term access to the material, but the three-month estimate provides little evidence that the advantage persisted. The present design therefore supports an immediate performance benefit more strongly than durable retention.

294. Mini-case 2 — Workplace | Productivity association

Finding:

Remote-work days correlate with higher individual output.

Weak:

Remote work increases productivity.

Calibrated:

Remote-work frequency was associated with higher individual output, but employees with greater autonomy were also more likely to work remotely. Selection into remote work therefore remains a plausible explanation for part of the association.

295. Mini-case 3 — AI benchmark | Small benchmark gain

Finding:

Model A outperforms Model B by 0.8 percentage points on one benchmark.

Weak:

Model A is superior.

Calibrated:

Model A achieved a small advantage on the tested benchmark. Whether this difference persists across domains, distribution shifts and repeated runs remains uncertain.

296. Mini-case 4 — Qualitative study | Divergent participant accounts

Finding:

Most learners value detailed feedback; a minority feel overwhelmed.

Weak:

Students prefer detailed feedback.

Calibrated:

Detailed feedback was experienced as supportive by many participants, but accounts of overload indicate that usefulness depended on learners’ existing revision routines and capacity to prioritise comments.

297. Mini-case 5 — Health-style observational example | Association with outcome

Finding:

Higher activity level associated with lower symptom score.

Weak:

Exercise reduces symptoms.

Calibrated:

Higher activity was associated with lower symptom scores. Because healthier participants may also be more able to remain active, reverse causality and residual confounding limit causal interpretation.

298. Mini-case 6 — Engineering | Performance–energy trade-off

Finding:

New system improves throughput 12% but uses 20% more energy.

Weak:

The new system performs better.

Calibrated:

The design increased throughput at the cost of higher energy use. Whether this represents an overall improvement depends on the deployment objective and the relative value assigned to speed and energy efficiency.

299. Mini-case 7 — Survey | Confidence vs competence

Finding:

Training increases self-reported confidence; objective performance unchanged.

Weak:

Training improved capability.

Calibrated:

Training increased perceived confidence without a corresponding improvement on the objective performance measure. The intervention therefore appears to have changed self-evaluation more clearly than demonstrated capability.

300. Mini-case 8 — Systematic review | Heterogeneous effects

Finding:

Pooled effect positive; studies vary widely.

Weak:

The intervention works.

Calibrated:

The pooled estimate favours the intervention, but substantial between-study variation indicates that the average effect does not describe all populations and implementation conditions equally well.

301. Mini-case 9 — Diagnostic test | High sensitivity, modest specificity

Weak:

The test is accurate.

Calibrated:

The test identified most positive cases but also produced a substantial number of false positives. Its usefulness therefore depends on whether missing cases or over-referral carries the greater cost in the intended setting.

302. Mini-case 10 — Longitudinal learning | Improvement with high attrition

Weak:

Scores improved steadily over the year.

Calibrated:

Scores among retained participants increased across the year, but attrition was concentrated among lower-performing learners. The observed trajectory may therefore overstate improvement in the original cohort.

303. Mini-case 11 — Corpus study | Frequency difference

Finding:

Phrase X more frequent in expert writing than student writing.

Weak:

Experts use X because it is better academic English.

Calibrated:

Phrase X was more frequent in the expert corpus. The difference may reflect disciplinary convention, genre distribution or greater control of stance; frequency alone does not establish that the phrase is intrinsically superior.

304. Mini-case 12 — A/B test | Click improvement

Finding:

New interface increases clicks but reduces task completion.

Weak:

The new design improves engagement.

Calibrated:

The interface increased click activity but reduced completed tasks, indicating that click-through and successful task completion moved in opposite directions. The data therefore do not support treating click rate alone as improved engagement.

305. Mini-case 13 — Policy pilot | Small positive effect, high implementation burden

Weak:

The policy should be expanded.

Calibrated:

The pilot produced a modest benefit, but implementation required substantially more staff time than routine practice. Expansion therefore depends on whether the benefit justifies the additional resource burden and whether the effect can be maintained under less intensive delivery.

306. Mini-case 14 — Non-inferiority | Correct interpretation matters

If a study is designed with a non-inferiority margin and meets the criterion:

The new treatment met the prespecified non-inferiority criterion relative to the comparator.

Do not automatically rewrite this as:

The treatments are identical.

307. Mini-case 15 — Failed replication | Evidence update

Original study large positive effect; replication near zero.

Weak:

The original study was wrong.

Calibrated:

The replication did not reproduce the original effect under the tested conditions, reducing confidence that the initial estimate generalises broadly. Differences in implementation, population or sampling error remain possible explanations for the discrepancy.

308. Mini-case 16 — Theory test | Both models predict same result

Weak:

The finding supports Theory A.

Calibrated:

The finding is compatible with Theory A, but because Theory B makes the same prediction, the result does not discriminate between the two accounts.

309. Mini-case 17 — Ceiling effect | Advanced learners

Finding:

No sentence-accuracy gain among very advanced learners.

Weak:

The intervention does not affect grammar.

Calibrated:

Sentence accuracy changed little in this advanced sample, although high baseline scores left limited room for detectable improvement. A less constrained measure would be needed to distinguish true absence of effect from a ceiling effect.

310. Mini-case 18 — Measurement mismatch | Engagement score

Finding:

Attendance rises, but attention not measured.

Weak:

Engagement improved.

Calibrated:

Attendance increased, indicating greater behavioural participation. The study did not measure cognitive or emotional engagement, so the broader construct remains unresolved.

311. Mini-case 19 — Short intervention | Novelty effect

Finding:

First-week enthusiasm high.

Weak:

Users strongly prefer the new system.

Calibrated:

Initial satisfaction was high during the first week, but the short observation window cannot distinguish durable preference from a novelty response.

312. Mini-case 20 — Multiple outcomes | One positive of ten

Weak:

The intervention improved creativity.

Calibrated:

One of ten secondary creativity measures favoured the intervention. Given the number of comparisons and the absence of a primary creativity effect, the isolated finding should be treated as exploratory.

313. Mini-case 21 — Mediation | Indirect-effect claim

Weak:

Motivation mediated the treatment effect.

Calibrated:

The estimated indirect effect through motivation was compatible with mediation under the model assumptions, but unmeasured mediator–outcome confounding remains a limitation.

314. Mini-case 22 — Qualitative saturation | Cautious reporting

Weak:

No new themes can exist.

Calibrated:

Additional interviews in the final sampling phase did not generate substantively new themes within the study’s sampling frame.

315. Mini-case 23 — Model robustness | Multiple specifications

Weak:

The result is robust.

Calibrated:

The direction and approximate magnitude of the association were similar across three prespecified model specifications, although all relied on the same observational dataset.

316. Mini-case 24 — External validation | Predictive model

Finding:

High internal test score; lower external-site score.

Calibrated:

Performance remained above baseline in the external dataset but declined relative to the internal test set, indicating partial rather than complete transportability across sites.

317. Mini-case 25 — Harms vs benefit | Decision balance

Finding:

Benefit improves 5%; adverse events rise 3%.

Weak:

The intervention is beneficial.

Calibrated:

The intervention improved the target outcome but also increased adverse events. Its overall value therefore depends on the relative importance, severity and reversibility of these competing outcomes.

318. Sentence transformation laboratory 1 | Association → causal overclaim → repair

Observed: Students who revise more often score higher.

Overclaim: Frequent revision causes higher scores.

Repair: Revision frequency was associated with higher scores; motivation and prior attainment remain plausible common causes.

319. Sentence transformation laboratory 2 | Null → equivalence error → repair

Observed: Difference = 1.2, 95% CI −4.0 to 6.4.

Overclaim: The methods are equally effective.

Repair: The study did not detect a clear difference, but the interval remains too wide to establish equivalence.

320. Sentence transformation laboratory 3 | Proxy → construct inflation → repair

Observed: Confidence score rises.

Overclaim: Competence improved.

Repair: Participants reported greater confidence; objective competence was not directly established by this measure.

321. Sentence transformation laboratory 4 | Single site → universal scope → repair

Overclaim: The method works for advanced learners.

Repair: The method improved the measured outcome among advanced learners in this high-support programme; replication is needed before broader generalisation.

322. Sentence transformation laboratory 5 | Short term → durable claim → repair

Overclaim: The intervention builds lasting knowledge.

Repair: The intervention improved one-week performance; durability beyond the measured follow-up remains unknown.

323. Sentence transformation laboratory 6 | Mechanism story → evidence-calibrated mechanism

Overclaim: The prompts worked because they reduced cognitive load.

Repair: Reduced cognitive load is one plausible explanation, but load was not measured and alternative mechanisms remain possible.

324. Sentence transformation laboratory 7 | Secondary positive → spin repair

Overclaim: The intervention improved learning.

Repair: The primary performance outcome did not improve clearly, although self-reported confidence increased as a secondary outcome.

325. Sentence transformation laboratory 8 | Subgroup mining → exploratory repair

Overclaim: The intervention works for beginners.

Repair: A post-hoc subgroup pattern suggested a larger effect among lower-baseline learners; the possible interaction requires preregistered replication.

326. Sentence transformation laboratory 9 | Review consensus inflation → repair

Overclaim: Research proves the intervention is effective.

Repair: Most included studies report positive short-term effects, but heterogeneity and risk-of-bias concerns reduce confidence in a single general effect estimate.

327. Sentence transformation laboratory 10 | “Should” leap → decision-aware repair

Overclaim: Schools should adopt the programme.

Repair: The findings justify further implementation testing; routine adoption should depend on replication, resource requirements and evidence of durable benefit.

328. Peer-review simulation | Reviewer 1: “You overstate mechanism”

Original:

The improvement occurred because the structured prompts increased metacognition.

Revised manuscript:

The improvement may reflect more explicit revision planning. Metacognition was not measured directly, and the structured condition differed from the comparator in several planning demands.

Response:

We agree that the original sentence overstated the mechanism. We have removed the causal metacognition claim and now present planning as one plausible explanation while identifying the component-confounding limitation.

329. Peer-review simulation | Reviewer 2: “Your limitations do not affect your conclusion”

Original limitation:

This study was limited by its small sample and single site.

Original conclusion:

The intervention is effective for advanced learners.

Revised:

The trial provides evidence of a near-term effect in this programme, but the single high-support site narrows transportability and the delayed estimate remains imprecise. The conclusion has therefore been restricted to the studied setting and time horizon.

330. Peer-review simulation | Reviewer 3: “Primary outcome is null; abstract highlights secondary outcome”

Correct repair:

  1. restore the primary null result to the abstract;
  2. label the positive secondary outcome as secondary;
  3. revise the Discussion hierarchy;
  4. revise the title if it implies overall efficacy.

This is not cosmetic revision. It changes the study narrative to match the prespecified evidence hierarchy.

331. Peer-review simulation | Reviewer 4: “Generalisation beyond sample is unsupported”

Original:

These findings show that bilingual learners benefit from structured feedback.

Revised:

These findings show a near-term benefit among the advanced bilingual volunteers studied. Effects in younger learners, beginners and compulsory-learning settings remain uncertain.

332. Peer-review simulation | Reviewer 5: “Discussion repeats Results”

Delete:

every repeated group mean, p-value and table cell that does not perform new inferential work.

Add:

  • relationship to prior evidence;
  • alternative explanations;
  • limitation consequences;
  • scope;
  • next discriminating test.

333. Discussion compression pass | Delete repetition before deleting reasoning

If the journal word limit requires a shorter Discussion, cut in this order:

  1. repeated numerical Results;
  2. duplicate literature examples;
  3. minor limitations;
  4. generic future-research language;
  5. decorative transitions.

Protect:

  • primary interpretation;
  • strongest rival;
  • most consequential limitation;
  • scope boundary;
  • final calibrated claim.

334. Discussion expansion pass | Add depth without padding

If a thesis/dissertation requires more depth, expand by adding:

  • stronger comparison of rival explanations;
  • methodological consequences of limitations;
  • boundary-condition analysis;
  • discipline-specific theoretical implications;
  • future studies that discriminate explanations;
  • integration across multiple findings.

Do not expand by repeating the same conclusion with new adjectives.

335. The final pre-submission Discussion checklist | Final gate

QuestionPass condition
Primary outcome first?Yes, clearly visible.
Null/negative findings discussed?Yes, proportionately.
Mechanism separated from effect?Yes.
Strongest rival included?Yes.
Literature relationship specific?Replicate/extend/qualify/contradict named.
Limitations alter claim?Each major limitation has consequence.
Generalisation bounded?Population/setting/time preserved.
Implication separate from recommendation?Yes.
Future work discriminating?Targets largest uncertainty.
Title/abstract aligned?No certainty or scope inflation.

336. The bilingual pre-submission checklist | 中文母语 learner final gate

Search your draft for English words likely to be literal translations of broad Chinese academic connectors:

  • prove;
  • clearly;
  • obviously;
  • therefore;
  • because;
  • should;
  • must;
  • significant;
  • to some extent;
  • it can be seen that.

For each, ask:

What exact evidential relation am I expressing?

Then rebuild from that relation rather than from the Chinese wording.

337. The one-sentence compression test | Can you state the paper accurately in one sentence?

Template:

In [population/context], [design/evidence] supports [main claim] over [time/outcome], while [major uncertainty] remains unresolved.

Example:

In advanced bilingual learners at one high-support programme, a randomised comparison supports a near-term independent-revision benefit from structured feedback, while mechanism, long-term durability and routine-setting transportability remain unresolved.

338. The two-sentence decision test | What should a decision-maker know?

Sentence 1:

What is established?

Sentence 2:

What uncertainty most limits action?

If the Discussion cannot answer these, it may be too diffuse.

339. The expert-opponent test | Would a sceptical expert recognise the evidence state?

Give the Discussion to someone who disagrees with your preferred interpretation. Ask them:

  • Did I state your strongest objection fairly?
  • Did I hide any inconvenient result?
  • Did I exaggerate certainty?
  • Did I misrepresent prior work?
  • Would you accept the final claim even if you reject my preferred mechanism?

340. The final principle | Discussion is disciplined permission

Every Discussion sentence asks for permission to move one step beyond the observed data.

The permission is earned by:

  • design;
  • precision;
  • replication;
  • measurement;
  • comparison with rivals;
  • consistency with the broader evidence;
  • transparent limitations;
  • scope control.

C1–C2 academic English is not the ability to make claims sound powerful. It is the ability to make powerful reasoning visible while keeping every claim inside the boundary the evidence can support.

高级学术英语真正难的,不是把句子写得“很有权威感”。真正难的是:让 reasoning 变得清楚,同时让每一个 claim 都停在 evidence 允许它停的位置。


Final quality gate | 最终质量门

Before publication, run this sequence one last time:

  1. Primary question → primary answer.
  2. Effect → mechanism separation.
  3. Preferred explanation → strongest rival.
  4. Prior literature → exact relationship.
  5. Null result → honest consequence.
  6. Limitation → claim reduction.
  7. Sample/setting/time → generalisation boundary.
  8. Implication → recommendation boundary.
  9. Future study → biggest unresolved uncertainty.
  10. Discussion → abstract/title consistency.
  11. English wording → Mandarin back-translation.
  12. Final claim → what would change your mind?

If all twelve connections remain visible, the Discussion is not merely fluent. It is structurally trustworthy.

Part VI — The 20,000-word transfer layer: full reconstruction, bilingual claim repair and mastery rubric | 第六部分:完整迁移训练

This final layer is deliberately practical. The earlier sections explain the intellectual machinery of Discussion; this section forces that machinery to operate repeatedly across unfamiliar claims, disciplines and bilingual formulations. The aim is not to memorise more phrases. The aim is to make accurate interpretation automatic enough that a Mandarin-speaking C1–C2 writer can enter a new research field and still control claim strength, evidence hierarchy and inferential boundaries.

341. Full reconstruction case B | 第二个完整案例:不是教育 feedback

All data below are fictional teaching material.

A workplace study examines whether a four-day workweek changes output and wellbeing. Twelve companies volunteer. Six adopt a four-day schedule for three months and six continue normal schedules. Assignment occurs at company level but is not random: firms choose whether to adopt the programme. Productivity is measured from completed project units, while wellbeing is measured using a self-report scale. At three months, adopter firms show an average 7% increase in project units and a 12% increase in wellbeing score. However, adopting firms had higher baseline employee autonomy and lower turnover before the programme began. Two adopter firms also introduced new project-management software during the study.

342. Start by separating the two outcomes | Productivity and wellbeing are not one result

A weak Discussion writes:

The four-day workweek improved employee performance and wellbeing.

This sentence collapses two outcomes and implies causality.

A stronger opening:

Companies adopting the four-day schedule showed higher three-month productivity and self-reported wellbeing than comparison firms, but the non-random adoption process limits causal attribution.

343. Identify the main causal threats | Selection + co-intervention

The strongest rival explanations are not mysterious. First, firms choosing the programme already differed in autonomy and turnover, so selection may explain part of the outcome gap. Second, two adopter firms introduced new software, creating a co-intervention. A serious Discussion must make these alternatives visible before offering a strong workplace-policy recommendation.

344. Write the productivity interpretation | Keep the outcome literal

The observed productivity difference is compatible with a benefit from the shorter workweek, but it is also compatible with pre-existing organisational differences and concurrent process changes. The present design therefore establishes an association between adoption and output more clearly than it establishes a causal schedule effect.

345. Write the wellbeing interpretation | Self-report boundary

Employees in adopter firms reported higher wellbeing, indicating a favourable perceived experience during the intervention period. Because wellbeing was self-reported and participants knew their company had adopted the new schedule, expectancy or social-desirability effects cannot be excluded.

346. Handle the baseline imbalance | Baseline difference is not a footnote

The Discussion should ask whether adjustment changes the estimate, but adjustment cannot guarantee removal of all selection. A calibrated sentence is:

Adjustment for baseline autonomy and turnover reduced but did not eliminate the between-group difference, although unmeasured organisational culture may still influence both programme adoption and subsequent performance.

347. Handle the software change | Co-intervention makes mechanism ambiguous

The introduction of new project-management software in two adopter firms further complicates attribution. If those firms contributed disproportionately to the productivity gain, the schedule effect may be overestimated.

This is a directional limitation: it identifies how the co-intervention could bias the estimate.

348. Generalisation from 12 volunteer firms | Do not jump to “companies”

The volunteer firms were already willing to redesign working arrangements and had relatively high baseline autonomy. The findings may therefore transport more readily to flexible knowledge-work organisations than to settings with fixed staffing, continuous service requirements or lower scheduling autonomy.

349. Practical implication without policy inflation | What can decision-makers do now?

The findings justify a larger controlled evaluation in organisations with different operational constraints. They are not sufficient to establish that a four-day schedule will increase productivity across industries.

350. Final calibrated conclusion for case B | Four layers

Voluntary adoption of a four-day schedule was associated with higher short-term productivity and reported wellbeing in this group of firms. Selection into adoption and concurrent software changes prevent a clean causal estimate, and the volunteer sample limits broader transportability. A cluster-randomised or stronger quasi-experimental study that holds other organisational changes constant would provide a more discriminating test. The present evidence supports further controlled implementation research rather than a universal productivity claim.

351. Why this second case matters | Transfer beyond familiar education examples

The reasoning system stays the same even though the subject changes. We still ask what was observed, which causal pathways remain open, which population was actually studied, whether an outcome is self-report or direct performance, and what next design would distinguish competing explanations. This is what transfer looks like.

352. Full reconstruction case C | AI-assisted writing system

All details below are fictional teaching material.

A study compares an AI-assisted drafting tool with a standard word processor among 180 university writers. Participants are randomised. The AI group produces essays rated 4.5 points higher immediately after a two-hour session. When both groups complete an unaided essay four weeks later, the difference is 0.8 points with a wide interval. AI users finish the assisted essay 18% faster but accept 14% of generated claims that independent fact-checkers classify as unsupported.

353. There are three outcomes and one trade-off | Quality, speed, transfer, factual risk

A weak narrative selects the attractive pair:

AI improves writing quality and speed.

A complete evidential state is more complicated:

  • assisted quality improved;
  • assisted completion time fell;
  • unaided four-week transfer was unclear;
  • unsupported-claim acceptance created a factual-quality risk.

354. Write the assisted-performance claim | Keep support condition visible

Access to the AI drafting tool improved essay ratings and reduced completion time during the supported writing session.

The word supported prevents assisted performance from being silently converted into unaided skill.

355. Write the transfer claim | Do not infer learning from tool-assisted output

The four-week unaided task showed little clear between-group separation, so the study provides stronger evidence for improved tool-assisted performance than for durable independent writing improvement.

356. Write the factual-risk interpretation | Benefit and harm can coexist

The productivity and quality gains were accompanied by a non-trivial rate of accepted unsupported claims. This trade-off indicates that higher rated output under assistance should not be treated as equivalent to more reliable writing.

357. Avoid one-dimensional “AI works” language | System evaluation is multi-objective

A mature Discussion can conclude that the tool is useful for speed and assisted performance while remaining uncertain about learning transfer and concerned about factual reliability. Academic precision permits mixed conclusions.

358. Mechanism hypotheses for the AI case | Several plausible stories

The immediate quality gain might reflect:

  • better idea generation;
  • more fluent language;
  • reduced drafting time;
  • better structure;
  • simple substitution of model output for learner reasoning.

The study must measure these processes before selecting one as the mechanism.

359. A discriminating next study | Remove the ambiguity

A useful next study could compare AI conditions that provide planning prompts only, language revision only and full text generation, while measuring process logs and delayed unaided writing. This design would test which component improves assisted output and whether any component produces transfer.

360. Final calibrated conclusion for case C | Do not collapse tool performance into learning

The AI tool improved immediate assisted essay performance and speed but did not establish durable unaided writing gains at four weeks. The factual-error acceptance rate further indicates that productivity benefits can coexist with reliability costs. Future work should separate AI components and measure delayed independent performance before tool-assisted quality is interpreted as learner development.

361. Full reconstruction case D | Systematic review with conflicting studies

Fictional example.

A review identifies 28 studies of vocabulary retrieval practice. Twenty studies report positive immediate effects, 11 include delayed tests, and only five follow learners beyond eight weeks. The pooled short-term effect is positive. Delayed effects are smaller and heterogeneous. Most studies involve novice learners; three advanced-learner studies show little average benefit.

362. Do not let the pooled number erase time and proficiency | Average effect has boundaries

Weak:

Retrieval practice is effective for vocabulary learning.

Stronger:

The review supports a short-term vocabulary benefit from retrieval practice, particularly among novice learners. Evidence for longer-term persistence is thinner and more heterogeneous, and the small advanced-learner evidence base does not establish a comparable average benefit.

363. Heterogeneity needs interpretation, not dismissal | “Studies were mixed” is too weak

Ask whether variation aligns with:

  • baseline proficiency;
  • feedback;
  • retrieval difficulty;
  • delay interval;
  • outcome design;
  • study quality.

Then distinguish prespecified moderator evidence from post-hoc explanations.

364. Thin evidence should sound thin | Five long-term studies are not 28 long-term studies

The overall study count must not create false confidence about a subgroup of the evidence. State the denominator relevant to each claim.

365. Final calibrated conclusion for case D | Evidence density matters

Retrieval practice has a comparatively well-supported short-term benefit, but evidence becomes less certain as follow-up length and learner proficiency increase. The next useful research priority is not another immediate novice test; it is adequately powered delayed work among more advanced learners using comparable retention outcomes.

366. Bilingual claim-repair laboratory: purpose | 不是词汇翻译,而是 inference translation

Each item below begins with a Chinese-style research note that could easily produce overstrong English if translated literally. Reconstruct the evidence relation first, then choose English.

367. Repair 1 — “研究证明该方法有效”

Evidence: one randomised short-term trial in one programme.

Risky:

The study proves that the method is effective.

Calibrated:

The trial provides evidence that the method improves the measured short-term outcome under the studied conditions.

368. Repair 2 — “数据说明学习能力提高了”

Evidence: one task score increased.

Calibrated:

The task score increased, but the study measured performance on this task rather than learning ability as a broader construct.

369. Repair 3 — “因此学生掌握了这个能力”

Evidence: immediate supported performance.

Calibrated:

The immediate task shows successful supported performance; independent mastery requires evidence from later unsupported transfer.

370. Repair 4 — “这可能是由于学生更有动力”

Motivation unmeasured.

Greater motivation is one possible explanation, but motivation was not measured and the present data cannot evaluate this mechanism directly.

371. Repair 5 — “两组没有显著差异,因此效果一样”

No statistically clear difference was detected, but the estimate remains too imprecise to establish equivalence.

372. Repair 6 — “明显改善”

Replace the vague booster with a magnitude:

Mean score increased by 5.6 points relative to the comparison condition.

373. Repair 7 — “具有长期作用”

Evidence ends at four weeks.

The benefit remained detectable at four weeks; persistence beyond that interval was not assessed.

374. Repair 8 — “适合所有学生”

Evidence: advanced adult volunteers.

The result applies directly to the advanced adult volunteers studied; effects in other learner groups require separate evidence.

375. Repair 9 — “值得广泛推广”

Evidence: one efficacy trial.

The result justifies broader implementation testing, but routine adoption requires evidence on feasibility, cost and effectiveness under ordinary conditions.

376. Repair 10 — “结果与理论一致,因此理论正确”

The finding is consistent with the theory but does not establish it uniquely because alternative accounts make similar predictions.

377. Repair 11 — “没有发现副作用”

Small study.

No adverse events were observed in this sample; the study is too small to exclude uncommon harms.

378. Repair 12 — “显著降低风险”

Give absolute and relative state where useful:

Risk fell from 10% to 7%, an absolute reduction of three percentage points.

Whether this is statistically or practically significant must be stated separately.

379. Repair 13 — “基本没有差别”

The estimated difference was small relative to the prespecified practical threshold, although uncertainty around the estimate should still be considered.

380. Repair 14 — “用户普遍认可”

Evidence: 18/24 interviewed users positive.

Eighteen of 24 interview participants described the system favourably; six raised concerns about workload or control.

381. Repair 15 — “进一步说明机制是……”

Mechanism evidence indirect.

The additional pattern strengthens the plausibility of the proposed mechanism but does not isolate it from competing explanations.

382. Repair 16 — “可以推断”

Ask whether it is deduction, statistical inference or speculation.

Possible reconstructions:

The estimate supports the inference that…

One plausible interpretation is…

The data do not permit a firm inference about…

383. Repair 17 — “普遍存在”

Evidence from one corpus/sample.

The pattern was common in the analysed corpus.

Do not universalise beyond the sampling frame.

384. Repair 18 — “进一步证实”

If independent replication:

The replication increases confidence in the direction of the effect.

If same dataset reanalysis:

The alternative analysis produced a similar estimate.

These are not the same evidential gain.

385. Repair 19 — “无法排除”

Residual confounding cannot be excluded.

Remember: this means the study cannot rule it out, not that confounding is probably the true explanation.

386. Repair 20 — “很可能”

Do not automatically choose very likely. Ask whether evidence supports a probability claim. Often:

The pattern is consistent with…

or:

This explanation is plausible because…

is more faithful.

387. Repair 21 — “效果不理想”

Replace evaluation with outcome:

The observed effect was smaller than the prespecified target difference.

388. Repair 22 — “结果令人满意”

Replace emotion with criterion:

The result met the prespecified accuracy and latency thresholds.

389. Repair 23 — “有一定局限性”

Do not translate as has certain limitations.

Name the limitation:

The absence of delayed follow-up limits conclusions about persistence.

390. Repair 24 — “在一定程度上支持”

State the dimension:

The result supports the predicted direction but leaves the magnitude uncertain.

391. Repair 25 — “可以认为”

Replace writer permission with evidence relation:

The evidence supports the interpretation that…

392. Repair 26 — “总体而言效果较好”

A multi-outcome study needs outcome-specific language:

The intervention improved speed and satisfaction, while accuracy changed little.

393. Repair 27 — “效果稳定”

Stable across what?

The estimate was similar across three prespecified model specifications and two independent samples.

394. Repair 28 — “趋势明显”

State slope/direction:

Scores increased at each assessment, with the largest rise between weeks two and four.

395. Repair 29 — “进一步验证”

Validation can mean replication, predictive validation, instrument validation or checking. Choose exactly:

The external dataset provided an independent test of predictive performance.

396. Repair 30 — “支持这一观点”

The finding adds evidence for this interpretation, although it does not distinguish it from the competing account.

397. Repair 31 — “有必要采取措施”

State decision basis:

Given the observed risk, low cost of intervention and reversibility of the proposed step, a limited preventive action is justified while uncertainty is reduced.

398. Repair 32 — “需要更多研究”

Name the missing evidence:

A longer follow-up with an attention-matched comparator is needed to distinguish transient performance gain from durable transfer.

399. Repair 33 — “具有启示意义”

Name the implication:

The pattern suggests that future feedback designs should test action clarity separately from feedback volume.

400. Repair 34 — “结论可靠”

Reliable in which sense?

The primary direction was consistent across sensitivity analyses, although external validity remains limited by the single-site sample.

401. Repair 35 — “差异主要来自”

If not decomposed experimentally:

The evidence is compatible with X contributing to the difference, but the study does not isolate X as the principal cause.

402. Repair 36 — “进一步表明”

Ask what changed in certainty:

The second sample reproduced the direction of the association, increasing confidence that the pattern is not unique to the first sample.

403. Repair 37 — “相互印证”

For mixed evidence:

The behavioural and interview findings converged on the same broad pattern, although both could still be influenced by the shared intervention context.

404. Repair 38 — “尚不能确定”

Name the unknown:

The study establishes the direction of the short-term effect but not its long-term magnitude.

405. Repair 39 — “有待进一步观察”

Replace passive vagueness:

A six-month follow-up is needed to determine whether the improvement persists after support is withdrawn.

406. Repair 40 — “研究具有较高价值”

Do not praise the study generically.

The study contributes a direct delayed-transfer test that previous supported-task designs did not provide.

407. Discussion mastery rubric | 100-point advanced scoring system

DimensionPointsMastery evidence
Primary-question fidelity10main result and conclusion preserve outcome hierarchy
Interpretation calibration15claims match design, magnitude, precision and time horizon
Alternative explanations10strong rivals are represented and compared fairly
Literature integration10replicate/extend/qualify/contradict relationships are explicit
Null/counterevidence honesty10inconvenient findings actively shape the final claim
Limitations and bias15each major limitation has an inferential consequence
Generalisability10population, setting, outcome and time boundaries remain visible
Implications/recommendations5decision language adds appropriate cost/risk/feasibility logic
Future research5next study discriminates among unresolved explanations
Bilingual epistemic precision10no certainty/causality inflation from literal Chinese transfer

408. What 90–100 looks like | C2-ready Discussion control

A 90–100 Discussion does not merely contain all expected headings. It makes the evidence hierarchy visible. Its strongest claim is easy to trace to the design; its weakest claim is explicitly marked as uncertain. Alternative explanations are not straw men. Limitations alter the conclusion rather than decorating it. Prior literature is organised by relationship. Mandarin-to-English translation does not silently increase certainty. A sceptical expert can disagree with the preferred interpretation while still accepting the fairness of the evidence map.

409. What 75–89 looks like | Strong C1 with local weaknesses

The main reasoning is sound, but one or two areas may drift: a mechanism is slightly too confident, a limitation lacks a clear consequence, a subgroup receives too much emphasis, or a recommendation moves faster than implementation evidence. The repair is local rather than architectural.

410. What 60–74 looks like | Fluent language, unstable inference

The prose may sound academic, but evidence and interpretation are not consistently separated. Common signals include repeated shows/proves, generic literature comparison, limitations that do not alter claims, missing null results, and broad generalisation. This is the level where learners often overestimate their readiness because grammar and vocabulary conceal reasoning weakness.

411. Below 60 | Discussion is still a persuasive essay rather than evidence-controlled research prose

The writer chooses a story first and uses findings to support it. Counterevidence is minimised, alternative explanations are absent, and causal or policy claims exceed the design. The repair is not more academic vocabulary. Return to the Discussion matrix and rebuild the inference chain.

412. Self-marking protocol | Do not give yourself style points for sounding advanced

For each rubric dimension, require evidence from the draft:

  • quote the sentence that states the primary answer;
  • quote the sentence that states the strongest rival;
  • quote the limitation-consequence pair;
  • quote the null-result sentence;
  • quote the scope boundary;
  • quote the final calibrated claim.

If you cannot point to the sentence, do not award the points.

413. Reverse outline the Discussion | Paragraph-by-paragraph map

After drafting, write one margin note for each paragraph:

P1 — primary answer.

P2 — mechanism + rival.

P3 — persistence + prior literature.

P4 — limitation + scope.

P5 — implication + next study.

If two paragraphs have the same note, merge or differentiate them.

414. Compression challenge | 2,000 words → 800 words → 250 words

Write a full 2,000-word Discussion. Then compress it to 800 words while preserving:

  • primary answer;
  • strongest rival;
  • most important literature relation;
  • most consequential limitation;
  • final scope;
  • next discriminating test.

Then compress the final claim to 250 words. If certainty or scope changes during compression, repair it.

415. Expansion challenge | 250 words → 1,500 words without repetition

Start from a calibrated abstract-style conclusion and expand only by adding new analytical jobs:

  1. interpretation;
  2. rival explanation;
  3. prior literature relationship;
  4. bias pathway;
  5. generalisation boundary;
  6. future discriminating study.

Do not expand by restating the result six times.

416. Spoken Discussion drill | Explain evidence aloud before writing

In two minutes, answer:

What happened?

What do you think it means?

What else could explain it?

What can you not claim?

What test would decide next?

Then write. Spoken reconstruction often reveals hidden overclaiming before formal language obscures it.

417. Bilingual oral drill | Mandarin explanation → evidence-state English

Explain the finding naturally in Mandarin first. Do not translate sentence by sentence. Instead extract five states:

  1. observation;
  2. certainty;
  3. alternative;
  4. boundary;
  5. next test.

Rebuild those states in English. This prevents Chinese discourse markers from controlling English epistemic force.

418. Peer teaching drill | Teach the difference between three sentences

Explain to another learner why these are different:

The result proves X.

The result supports X.

The result is consistent with X.

If you cannot explain what additional alternative explanations each sentence permits, you do not yet control the verbs.

419. Adversarial rewrite drill | Make the claim too strong, then repair it

Take a valid cautious sentence and deliberately overclaim it. Then identify every extra assumption introduced by the stronger version.

Example:

Valid: Higher tool use was associated with higher scores.

Overclaim: Using the tool causes students to learn more.

Extra assumptions:

  • causal direction;
  • no confounding;
  • score = learning;
  • effect applies to students broadly.

420. Reverse adversarial drill | Make the claim too weak, then restore earned confidence

Over-hedged:

There may perhaps be a possibility that the intervention could have had some effect.

If a randomised precise trial shows a clear effect:

The intervention improved the measured one-week outcome in this sample.

Academic caution is not permanent timidity.

421. Final mastery benchmark | Unseen paper, no model answer

Choose an unfamiliar empirical article from a credible source. Read only the Introduction, Methods and Results. Do not read the authors’ Discussion yet.

Create:

  1. primary evidence state;
  2. confidence map;
  3. two mechanisms;
  4. two rival explanations;
  5. three limitation-consequence pairs;
  6. one generalisation boundary;
  7. one practical implication;
  8. one discriminating future study;
  9. an 800–1,200 word Discussion;

Then read the published Discussion and compare. The goal is not to match the authors. The goal is to detect where either interpretation exceeds or understates the evidence.

422. Final series-level reason for this lesson | 为什么我们不是只教“Discussion 常用句”

A phrase list can make a paragraph look academic in five minutes. It cannot make the reasoning trustworthy. The purpose of this Advanced English Chinese Edition is to help Mandarin-speaking learners control English at the level where language and judgement become inseparable: the exact scope of a claim, the distance between evidence and inference, the difference between possibility and probability, the boundary between implication and recommendation, and the ability to remain precise when the evidence refuses to tell a simple story.

That is why this lesson is long. The target is not recognition. The target is transfer.

423. Exit standard | You are ready to move on when…

You can take an unfamiliar Results section and, without relying on memorised sentence frames:

  • state what is established;
  • separate effect from mechanism;
  • identify a serious rival explanation;
  • make a limitation change the claim;
  • bound generalisation;
  • write a proportionate implication;
  • design the next test;
  • reconstruct the same evidence state in Mandarin and English without certainty drift.

When these moves are reliable, Discussion writing becomes more than fluent English. It becomes disciplined knowledge work.


Next lesson remains reserved | 下一课继续

EDKS-ADV-ZH-0021 · Lesson No.021 · Write a Conclusion That Closes the Argument Without Making It Bigger | 把研究结论写到真正收束,而不是最后再扩大主张

The next lesson will separate Conclusion from Abstract and Discussion, show how to compress the final answer without certainty inflation, and train endings that state contribution, boundary and next implication without reopening the paper.