Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

How to Learn Advanced English (Chinese Edition) | Lesson No.024 | Write a Peer Review Report That Improves the Paper Without Rewriting It as Your Own | 第024课:写好 Peer Review:严格、具体、可执行,但不要把别人的论文改成自己的

Series ID: EDKS-ADV-ZH-0024 · How to Learn Advanced English (Chinese Edition) · Lesson No.024 · C1 → C2

Write a Peer Review Report That Improves the Paper Without Rewriting It as Your Own | 写好 Peer Review:严格、具体、可执行,但不要把别人的论文改成自己的

A good peer reviewer is not a hidden co-author. The reviewer’s job is not to redesign the paper until it matches the reviewer’s preferred theory, method, style or research programme. The job is to test whether the manuscript’s own question, design, evidence and conclusions hold together—and to tell the editor and authors exactly where they do or do not.

好的 peer review,不是“我会怎么写这篇论文”。Reviewer 的真正工作,是判断这篇论文按照它自己的 research question、design、evidence 与 claim,是否成立;哪里不成立;哪里可以修;哪些问题是 validity-critical;哪些只是 preference。

The governing sequence is:

EXPERTISE → CONFLICT/CONFIDENTIALITY → MANUSCRIPT IDENTITY → VALIDITY → EVIDENCE → CLAIMS → MAJOR/MINOR PRIORITY → ACTIONABLE COMMENT → EDITOR RECOMMENDATION.

Expertise → conflict/confidentiality → manuscript identity → validity → evidence → claims → major/minor priority → actionable comment → editor recommendation。


Part I — Before you review: expertise, ethics, scope and reviewer identity | 第一部分:真正开始读论文以前

1. Accept only if you can review the paper’s core scientific job

Major publishers and reviewer-ethics guidance begin with the same principle: accept a review only when you can evaluate the central scientific work to a professional standard. See the public reviewer guidance from Elsevier and COPE.

2. You do not need to be expert in every subfield | But disclose your limits

A paper may combine domain science, statistics, machine learning, qualitative methods, measurement and implementation. If one component lies outside your expertise, state that to the editor rather than bluffing competence. Nature’s reviewer guidance explicitly asks reviewers to indicate aspects they could not assess fully.

3. Reviewer expertise has boundaries | Do not bluff competence

Weak reviewer behaviour:

I am unfamiliar with causal inference, but the model seems fine.

Better:

I can assess the educational design and outcome interpretation, but the causal-identification strategy is outside my expertise and may benefit from additional statistical review.

4. Time is an ethical variable | Late review can damage the process

Only accept if you can meet the deadline or negotiate a realistic extension promptly. A rushed review can be as unhelpful as a late one.

5. Conflict of interest should be disclosed before reviewing

Potential conflicts can include recent collaboration, direct competition, financial interests, institutional relationships, personal relationships or strong prior involvement with the work. If a relationship may affect impartiality, notify the journal and ask the editor whether you should proceed. The COPE Ethical Guidelines for Peer Reviewers provide a useful baseline.

6. Intellectual disagreement is not automatically a conflict

You may disagree strongly with the authors’ theory and still review fairly. The test is whether you can assess the evidence without requiring the manuscript to adopt your preferred worldview.

7. If you recognise the authors in double-blind review | Consider whether that knowledge creates conflict

If suspected identity creates a possible conflict, tell the editor rather than silently continuing.

8. Confidentiality is not optional | Manuscript content is not yours

Unpublished manuscripts, figures, ideas, datasets and reviewer materials are confidential. Do not discuss them publicly or use them for personal advantage. This is a core principle in both COPE and Elsevier reviewer guidance.

9. Do not share the manuscript casually with students or colleagues

If you want to involve a trainee or co-reviewer, follow the journal’s permission and co-review policy first. The invited reviewer remains responsible for the review.

10. Do not use unpublished ideas from the manuscript in your own work

Confidential access is not an information advantage you are entitled to exploit.

11. Do not upload confidential manuscripts into external generative-AI systems unless the journal explicitly permits it

Reviewer policies can differ, so check the journal. Nature’s current reviewer guidance states that confidential manuscripts should not be uploaded into generative-AI tools, and publisher policies generally connect AI use with confidentiality and human accountability. See Nature | How to Write a Report and the relevant journal instructions.

12. Reviewer accountability remains human

If any tool helps you check a claim, you still own the judgement. Verify every criticism before submitting it.

13. Ethical concern is not something to investigate independently

If you suspect plagiarism, fabrication, image manipulation, duplicate submission or another integrity issue, raise it confidentially with the editor. Do not contact the authors, conduct a private investigation or circulate the manuscript. COPE’s reviewer guidelines are the appropriate public reference.

14. Scientific criticism and misconduct allegation are different

Scientific criticism: the analysis does not support the conclusion.

Integrity concern: data may have been fabricated or manipulated.

Do not escalate ordinary methodological disagreement into misconduct language.

15. Scope of review should be clear | What has the editor asked you to assess?

Some invitations request a full scientific review; others request statistical, technical, clinical, methods or revision-only assessment. Review within the requested remit and disclose what you did not evaluate.

16. Journal-specific reviewer instructions outrank generic habits

Always read the journal’s reviewer instructions. Public guidance from Elsevier, Springer Nature and other publishers provides useful defaults, but the target journal’s rules control the actual review.

17. Read the journal’s aims and article criteria | The paper is reviewed for a venue

A methodologically sound paper may still be out of scope for a particular journal. That is a venue-fit judgement, not proof that the research is poor.

18. Editor decides publication | Reviewer recommends

Your recommendation supports the editor; it does not replace the editorial decision. Write accordingly.

19. Do not write your report as though you are issuing the final verdict

Weak:

This paper must be rejected.

Better:

I recommend rejection because the primary causal claim cannot be evaluated from the cross-sectional design and would require a fundamentally different study.

20. First read: identify the paper’s intended research job

Before criticising details, answer:

  • What question is the manuscript asking?
  • What type of evidence does it use?
  • What does it claim to contribute?
  • What decision does it ask the reader to accept?

21. Summarise the paper in your own words before judging it

Reviewer guidance from Elsevier and Springer Nature recommends a short opening summary. A useful template is:

In [population/context], the authors use [design] to test/describe/interpret , report [main result], and conclude [main claim].

22. The summary is a diagnostic | If you cannot summarise the study, read again

The manuscript may be unclear—or you may not yet understand it. Do not write aggressive criticism until you can reconstruct the study fairly.

23. Distinguish the paper’s question from your preferred question

Authors ask:

Does X predict Y?

Reviewer wants:

What mechanism causes Y?

The second may be interesting without being necessary to validate the first.

24. Do not turn every paper into your research programme

You may wish the authors had another population, a different model, more measures, a mechanistic experiment, another theory or a longer follow-up. Ask whether the addition is necessary for validity or merely expands ambition.

25. Validity-critical vs ambition-expanding request | Core distinction

Validity-critical: without the change, the manuscript’s current conclusion is unsupported.

Ambition-expanding: the paper is valid within scope, but the reviewer wants a broader or deeper paper.

26. Example of validity-critical request

Authors claim a causal treatment effect but the primary model ignores cluster randomisation. A cluster-aware analysis may be necessary before the central effect can be evaluated.

27. Example of ambition-expanding request

Authors show an internally valid adult effect; reviewer asks for a secondary-school sample. Unless the manuscript already claims children, that may belong to future research.

28. Reviewer scope creep | The hidden problem in “please also…”

Ten individually reasonable “please also” requests can transform one focused paper into several new studies.

29. The minimal-sufficient-review principle

Request the smallest set of changes needed to make the paper’s own contribution valid, interpretable and fairly scoped.

30. Peer review is not copyediting | Science before commas

Reviewer guidance from Nature Protocols and Springer Nature emphasises scientific assessment rather than line-by-line grammar correction.

31. Language matters when meaning becomes unclear | Distinguish language from science

Useful comment:

The description of the primary outcome is difficult to follow, and I could not determine whether the same score is used in Methods and Results. Please define the outcome once and use a stable term throughout.

32. Unhelpful language policing

The authors should use more elegant English.

This is neither specific nor scientifically actionable.

33. Reviewer should identify strengths too | Not as politeness theatre

Scientific strengths help editors judge the paper and help authors know what should be preserved during revision. Springer Nature’s public reviewer guidance explicitly encourages constructive identification of strengths and weaknesses.

34. Strengths can be scientific

  • strong design;
  • useful dataset;
  • clear preregistration;
  • important negative result;
  • external validation;
  • transparent limitations;
  • rare population;
  • good replication.

35. Do not invent praise to soften a rejection

Be fair, not performative.

36. Separate “I dislike this” from “this is invalid”

Preferences can involve theory, software, terminology, figure style or statistical framework. Preference is not error.

37. Distinguish evidence-based criticism from reviewer opinion

Evidence-based: The causal conclusion exceeds the design because exposure and outcome were measured cross-sectionally and temporal order is not established.

Preference: I would have preferred a Bayesian analysis.

If the current analysis is valid, the preference should not be treated as a major flaw.

38. A reviewer comment should identify the threatened inference

Do not merely announce a technique. Explain what becomes unreliable and why.

39. Preference-based comments should be optional where possible

Label them clearly so authors can distinguish requirements from presentation suggestions.

40. Reviewer should not demand citations for visibility

Elsevier’s public reviewer guidance warns against suggesting citations to the reviewer’s or associates’ work unless there is a genuine scientific reason.

41. Citation suggestion should answer a gap | Not increase your citation count

Good:

The Discussion states that no delayed-transfer studies exist. Study X directly tests delayed transfer and should be considered because it changes the novelty claim.

42. Poor citation request

Please cite my three papers on this topic.

43. Reviewer should not speculate about author motives

Weak:

The authors clearly chose this analysis to obtain significance.

Better:

The manuscript does not explain why this model was selected over the prespecified alternative, and the choice materially changes the result. Please clarify the analysis-selection rule.

44. Critique the manuscript, not the authors

Weak:

The authors do not understand statistics.

Better:

The interpretation of the interaction is not supported by the reported model because significance in one subgroup and non-significance in another do not establish a subgroup difference.

45. Ad hominem criticism is inappropriate

COPE, Elsevier and Springer Nature all emphasise objective, respectful review. A reviewer can be severe about a claim without attacking the people who wrote it.

46. Reviewer bias can enter before the Methods

Be alert to reactions based on institution prestige, country, author identity, topic politics, disciplinary status, language fluency or theory allegiance.

47. Review the work, not the status identity of the authors

The COPE reviewer guidelines are a useful reminder that origin and author characteristics should not distort scientific judgement.

48. Poor English is not poor science | But unclear English can obstruct assessment

If language prevents reliable assessment, tell the editor. If the science is understandable, do not downgrade it because the authors are non-native writers.

49. Reviewer recommendation should follow the report | Not lead it

Do not decide “reject” in minute five and then search for supporting reasons. Assess validity, contribution and repairability first.

50. Recommendation categories differ by journal | Use journal definitions

Common categories include accept, minor revision, major revision and reject. Exact meanings vary.

51. Recommendation is not a score for author quality

It is a recommendation about this manuscript in relation to journal criteria and revision feasibility.

52. Major revision should mean major scientific work, not many commas

Examples include new analysis, major construct correction, claim reframing, a critical control, substantial Methods transparency or a genuinely necessary additional experiment.

53. Minor revision should not hide a validity-critical issue

If one flaw could overturn the primary conclusion, the review is not minor simply because the comment is short.

54. Reject can be constructive | Explain the non-repairable core issue

Nature’s reviewer guidance encourages clear explanation of weaknesses, including in negative reviews. A rejection report should tell authors what cannot be repaired within the present study.

55. Reject vs major revision | Repairability test

Ask:

  • Can existing data fix the problem?
  • Can a feasible analysis fix it?
  • Can narrower claims fix it?
  • Would fixing it require a fundamentally new study?

56. Major revision should not mean “reject, but first do six months of new research”

If the paper requires a new research programme before its core claim can be tested, tell the editor clearly.

57. Confidential comments to the editor are not a place to hide scientific criticism

Public guidance from Nature and Nature Protocols stresses that substantive scientific evaluation should be communicated to authors; confidential comments are for genuinely sensitive/editorial matters.

58. What belongs confidentially with the editor?

  • suspected misconduct;
  • conflict concerns;
  • sensitive prior knowledge;
  • identity/confidentiality issues;
  • editorial fit/recommendation context;
  • need for specialist review.

59. What should not be hidden from authors?

If a scientific weakness contributes to a major-revision or rejection recommendation, authors should normally see and be able to respond to that scientific concern.

60. Part I operating rule | Reviewer discipline begins before critique

Before judging a manuscript, control yourself: your expertise, conflicts, confidentiality, theory preferences, scope ambitions and language expectations. A reviewer who cannot separate the paper’s problem from the reviewer’s preferences cannot produce a fair report.

Reviewer 真正的第一项能力,不是挑错,而是先控制自己的 bias、scope creep 与 preference。只有这样,后面的“严格”才是公平的严格。


Part I establishes reviewer ethics and scope. Part II builds the validity-first reading system: how to move from research question to design, data, analysis, results, interpretation and conclusions before deciding what counts as a major or minor concern.

Part II — Read for validity before style | 第二部分:先判断研究是否站得住,再讨论写得漂不漂亮

61. The reviewer’s first scientific question | Does the design answer the stated question?

Start with the match between:

research question → design → evidence → claim.

If these four objects do not align, the manuscript has a structural problem.

62. Major concern: question–design mismatch

Example:

Question claims causal effect.

Design is cross-sectional association.

This is major because the central inference cannot be repaired by style editing.

63. Minor concern: unclear wording of a valid question

If the design answers the right question but the wording is vague, that is usually repairable.

64. Preference: reviewer would have asked a different question

This is not a flaw unless the authors’ chosen question is incoherent, trivial for the venue or unsupported by the design.

65. Build a validity map before writing comments

ObjectReviewer question
QuestionWhat exactly is being tested/described/interpreted?
DesignCan this design answer that question?
SampleWho/what does the evidence represent?
MeasureDoes the measure represent the claimed construct?
AnalysisDoes the analysis estimate the relevant quantity?
ResultWhat does the evidence actually show?
DiscussionDoes interpretation exceed result/design?
ConclusionDoes final claim preserve boundaries?

66. Research question type matters | Different questions require different evidence

Common types:

  • descriptive;
  • associational;
  • causal;
  • predictive;
  • diagnostic;
  • mechanistic;
  • interpretive;
  • qualitative experiential;
  • feasibility;
  • implementation;
  • review/synthesis.

67. Descriptive question | Do not demand causal design

If authors ask:

How common is X in population Y?

The key concerns are sampling, measurement and denominator—not randomisation.

68. Associational question | Confounding matters, but the paper may remain descriptive

Do not force causal language onto an associational paper simply because causality is more exciting.

69. Causal question | Identification strategy is central

Ask:

  • Is exposure assigned or observed?
  • What confounders exist?
  • Is temporal order clear?
  • Is there a valid comparison?
  • Are post-treatment variables controlled incorrectly?
  • Could selection bias explain the effect?

70. Predictive question | Accuracy is not causality

Ask about:

  • train/validation/test separation;
  • external validation;
  • calibration;
  • discrimination;
  • distribution shift;
  • data leakage;
  • baseline comparison.

71. Diagnostic question | Thresholds and population prevalence matter

Do not collapse sensitivity, specificity and predictive value into “accuracy.”

72. Mechanistic question | Mechanism requires mechanism evidence

A treatment effect does not automatically establish how the treatment works.

73. Qualitative question | Do not demand population prevalence from meaning-focused sampling

Ask whether data generation, sampling logic and analysis support the themes/interpretations claimed.

74. Feasibility question | Do not demand efficacy from a feasibility study

If the study is powered for recruitment, retention and delivery, outcome trends may remain exploratory.

75. Systematic-review question | Eligibility criteria must match the question

If the review claims to cover “all learners” but includes adults only, scope is incoherent.

76. Introduction review | Does the gap exist?

Check:

  • is prior evidence represented fairly?
  • are contradictory studies omitted?
  • is novelty exaggerated?
  • does the gap lead naturally to the actual study?

77. Major Introduction concern | False gap drives false contribution

If earlier literature already answers the claimed “unknown,” the manuscript’s contribution may need reframing.

78. Minor Introduction concern | Excessive background

Can often be fixed by compression.

79. Preference | Reviewer wants their favourite theory added

Add only if the theory genuinely changes interpretation or fairness.

80. Methods review | Can another expert understand what was actually done?

Nature and Elsevier both emphasise methods, data quality and statistical appropriateness as central reviewer jobs.

81. Population/sample validity | Who entered the study?

Ask:

  • eligibility;
  • recruitment;
  • sampling method;
  • attrition;
  • exclusions;
  • analysis population.

82. Major sample concern | Selection destroys target-population claim

Example:

Authors generalise to all patients but sample only volunteers from a specialist clinic.

83. Minor sample concern | Flow diagram unclear

If the underlying counts are consistent but presentation is confusing, request clearer flow.

84. Preference | Reviewer wishes sample were larger

“Larger is better” is not enough. State what inference is imprecise or invalid at current n.

85. Sample-size criticism should name consequence

Possibilities:

  • wide intervals;
  • unstable subgroup estimates;
  • model overfitting;
  • rare-event uncertainty;
  • inability to test interaction;
  • limited external validity.

86. Intervention/exposure validity | What exactly is X?

If an intervention bundles:

  • more contact time;
  • more feedback;
  • new prompts;
  • new teacher;
  • new software;

then attributing the effect to one component may be unjustified.

87. Major intervention concern | Comparator makes effect uninterpretable

Example:

treatment group receives both new method and twice the instructional time.

88. Minor intervention concern | Description insufficient for replication

Request concrete details.

89. Measure validity | Does the operational measure match the construct?

Examples of possible drift:

  • attendance → engagement;
  • confidence → competence;
  • clicks → successful task completion;
  • benchmark score → general capability;
  • self-report → objective behaviour.

90. Major measurement concern | The measure cannot support the central construct

This may require claim narrowing or different data.

91. Minor measurement concern | Reliability/scale description missing

Can often be fixed by reporting details.

92. Reviewer should distinguish measure limitation from author misconduct

A weak instrument is a scientific limitation unless evidence suggests something more serious.

93. Timing validity | When were exposure and outcome measured?

Temporal order is critical for causal interpretation and durability claims.

94. Major timing concern | “Long-term” conclusion from immediate post-test

Request exact time language or additional follow-up if long-term claim is essential.

95. Minor timing concern | Time labels inconsistent

“Post-test,” “T2” and “week 1” may refer to same assessment; ask authors to standardise.

96. Analysis review | Does the model estimate the question?

Do not ask whether the analysis looks sophisticated. Ask whether it is appropriate.

97. Major analysis concern | Wrong unit of analysis

Clustered data treated as independent can understate uncertainty.

98. Major analysis concern | Post-treatment adjustment

Conditioning on variables caused by treatment can distort causal estimates.

99. Major analysis concern | Leakage in predictive modelling

If test information leaks into training, reported performance may be invalid.

100. Major analysis concern | Outcome switching or cherry-picked model

Compare protocol/preregistration where applicable.

101. Minor analysis concern | Missing effect sizes/intervals

Often repairable without changing core analysis.

102. Preference | Reviewer prefers different software

R vs Python vs Stata is not a scientific issue if the implementation is correct.

103. Preference | Reviewer prefers Bayesian or frequentist framework

Do not demand a framework swap unless the current analysis fails the research question or field/journal requirement.

104. Statistical significance | Reviewer should not equate p with importance

Ask about:

  • effect size;
  • precision;
  • practical meaning;
  • multiplicity;
  • model assumptions.

105. Null result | Do not ask authors to claim “no effect” automatically

Check confidence interval/equivalence logic.

106. Subgroup claims | Interaction, not separate significance

A significant effect in Group A and non-significant effect in Group B does not prove groups differ.

107. Mediation claims | Require temporal/causal assumptions

Do not treat simple correlation among X, M and Y as mechanism proof.

108. Multiple testing | Does the manuscript search for significance?

Look for:

  • many outcomes;
  • many subgroups;
  • many models;
  • one highlighted positive result.

109. Robustness | Ask whether reasonable analytic choices change conclusion

Sensitivity analyses are useful when they address a real assumption—not as ritual extra work.

110. Results review | Are the data reported before interpretation?

Look for:

  • sample flow;
  • primary outcome first;
  • effect estimate;
  • uncertainty;
  • null/negative findings;
  • secondary/exploratory labels.

111. Major Results concern | Primary outcome hidden

If the main outcome is null but secondary positive findings dominate, hierarchy may be distorted.

112. Major Results concern | Numbers do not reconcile

n in text, table and figure differ with no explanation.

113. Major Results concern | Figure contradicts prose

Ask authors to reconcile underlying data, not merely redraw.

114. Minor Results concern | Repetition of table values

Can be improved for readability.

115. Minor Results concern | Figure caption does not define error bars

Nature specifically asks reviewers to assess statistical tests and definition of error bars/uncertainty.

116. Discussion review | Does interpretation match evidence?

This is where many papers overclaim.

117. Major Discussion concern | Association becomes causation

118. Major Discussion concern | Measured proxy becomes broader construct

119. Major Discussion concern | Mechanism invented

120. Major Discussion concern | Primary null result disappears

121. Major Discussion concern | Limitation acknowledged but ignored

If authors admit severe selection bias but still generalise broadly, the limitation has not changed the claim.

122. Minor Discussion concern | Literature comparison too generic

“Consistent with prior studies” may need specifics.

123. Preference | Reviewer wants authors to adopt reviewer’s theory

Ask whether current theory is incompatible with evidence. If not, alternative theory can be suggested as comparison rather than replacement.

124. Conclusion review | Does the final claim preserve the paper’s evidence budget?

Check:

  • causality;
  • population;
  • time;
  • outcome;
  • mechanism;
  • recommendation strength.

125. Major Conclusion concern | Policy recommendation exceeds study

A small efficacy study may not justify system-wide adoption.

126. Major Conclusion concern | “Durable” without durable evidence

127. Minor Conclusion concern | Generic future work

Ask for a more specific unresolved question if useful.

128. Tables and figures are part of the scientific claim

Review:

  • axis labels;
  • units;
  • denominators;
  • uncertainty;
  • colour/legend;
  • sample identity;
  • caption interpretation.

129. Supplementary material is also review material where expected

Nature asks reviewers to review all data including Supplementary Information.

130. Supplement cannot hide a result that changes main conclusion

Major sensitivity reverses effect? Main text must reflect it.

131. Reporting guidelines | Use them as completeness tools

For relevant designs, consider CONSORT, STROBE, PRISMA, TRIPOD, COREQ/SRQR or other field-specific standards.

132. Do not use a checklist as a substitute for thinking

A manuscript can satisfy reporting items and still have a flawed design.

133. Ethics approval | Check appropriateness, not just presence

Human/animal studies may require ethics approval, consent or trial registration depending on context.

134. Data availability | Does claim depend on inaccessible evidence?

Availability requirements vary by journal and data type. If restrictions exist, authors should explain them.

135. Reproducibility | What can another researcher reconstruct?

Reviewers should ask whether Methods, code, parameters and data provenance are sufficient for the field’s norms.

136. AI/ML reproducibility | Versioning matters

Check:

  • model version;
  • prompt/procedure;
  • dataset split;
  • random seeds;
  • evaluation metric;
  • external validation;
  • contamination risk.

137. Qualitative validity | Use genre-appropriate criteria

Consider:

  • sampling logic;
  • data saturation/information power where relevant;
  • reflexivity;
  • analytic transparency;
  • negative/deviant cases;
  • traceability from data to interpretation;
  • transferability.

138. Do not impose quantitative standards onto qualitative work

“Why is n not powered?” may be meaningless for the qualitative research job.

139. Systematic review validity | Search and selection are methods

Check:

  • question/eligibility match;
  • search completeness;
  • selection process;
  • risk of bias;
  • synthesis method;
  • heterogeneity;
  • certainty.

140. Meta-analysis | Pooled effect cannot erase heterogeneity

Review whether subgroup/meta-regression interpretations are prespecified and adequately supported.

141. Engineering validity | Operating envelope matters

Prototype performance under ideal conditions should not become deployment robustness without testing.

142. Humanities/interpretive validity | Evidence and scope still matter

Do not demand randomisation; ask whether textual/historical evidence supports the interpretation and whether corpus/archive boundaries are respected.

143. Replication study | Novelty criteria differ

A high-quality replication can be valuable even without a new theory. Do not penalise it for not being “novel” in the same way as an exploratory discovery paper.

144. Negative result | Can be important

Do not treat null findings as failed papers if design and question are strong.

145. Reviewer’s validity ladder | Rank problems by what they threaten

  1. Fatal/non-repairable: central claim cannot be supported by available design/data.
  2. Major repairable: analysis/interpretation/measurement can be corrected with existing or feasible evidence.
  3. Minor: clarity/reporting/local interpretation.
  4. Preference: alternative but not necessary choice.

146. Fatal concern should be explained in one clear causal chain

The manuscript claims X. The design only observes Y. Because Z is unmeasured/undefined, X cannot be estimated from these data. A different study would be needed.

147. Major repairable concern should include a repair target

Account for clustering in the primary uncertainty estimate and update Abstract/Discussion if the interval materially changes.

148. Minor concern should not be written like a crisis

Please define the error bars in Figure 2.

149. Preference should be labelled as optional

Optional: Figure 3 may be easier to read as two panels, but the current presentation is scientifically interpretable.

150. Major vs minor is about scientific consequence, not comment length

A one-line issue can be fatal. A page of style suggestions can be minor.

151. Do not inflate minor issues to justify a recommendation

Your decision should emerge from validity and journal criteria, not from comment volume.

152. Do not bury the most important issue at comment 17

Order major comments by scientific importance.

153. Major comment 1 should often identify the central threat

Make the report navigable.

154. If there are no major scientific problems | Say so

Do not invent major comments because a review feels incomplete without them.

155. Minor comments can still materially improve clarity

But they should not overshadow the validity assessment.

156. The reviewer’s one-page evidence map | Before writing report

LayerYour note
QuestionWhat is being asked?
DesignCan it answer?
Primary evidenceWhat directly answers?
Main uncertaintyWhat could change interpretation?
ScopeWho/where/when?
ContributionWhat is genuinely added?

157. If you cannot fill this map, read again before writing

158. The “authors could fix this how?” test | Every major comment should have a response path

Possible response paths:

  • new analysis;
  • claim narrowing;
  • clarification;
  • new control;
  • additional data;
  • better uncertainty reporting;
  • literature correction;
  • restructuring.

159. If no feasible response path exists | Tell editor clearly

This may support rejection rather than an impossible major revision.

160. Part II operating rule | Review validity before preference

A reviewer earns the right to request changes only after showing what inference is threatened, why it is threatened, and whether the threat is fatal, major, minor or optional.

好的 reviewer comment 不是“我建议你改”,而是“这个问题威胁哪一个 inference、为什么、严重到什么程度、最小的有效 repair 是什么”。


Part II builds the validity-first reading system. Part III will convert those diagnoses into reviewer comments that are specific, proportionate and actionable without becoming commands to rewrite the authors’ study as the reviewer’s own.

Part III — Write actionable comments without taking ownership of the paper | 第三部分:写得具体,但不要把作者变成执行你研究计划的人

161. The anatomy of a strong reviewer comment | Four parts

A strong comment usually contains:

  1. Observation: what you see.
  2. Consequence: why it matters scientifically.
  3. Repair target: what must become true.
  4. Implementation freedom: avoid dictating one unnecessary solution.

162. Weak comment: “Methods are unclear.”

This does not tell authors what cannot be understood.

163. Strong repair

The primary outcome is described as “independent revision” in the Introduction but the Methods do not state whether feedback remained visible during the task. This distinction determines whether the outcome measures supported or unsupported performance. Please define the task conditions explicitly and align the outcome terminology across sections.

164. Weak comment: “Use a better statistical model.”

Better relative to what problem?

165. Strong repair

Participants are nested within six instructors, but the primary model treats all observations as independent. Please account for instructor-level clustering or explain why clustering is negligible; the uncertainty estimate and any downstream claims should be updated if the interval changes materially.

166. Weak comment: “Run propensity-score matching.”

This dictates a method before defining the inferential problem.

167. Strong repair target

Because treatment groups differ substantially in baseline autonomy, the current unadjusted comparison may be confounded. Please address baseline imbalance using an analysis justified by the design and causal structure, and report how the adjustment affects the estimate.

168. The reviewer should name the required property, not always the required technique

Required property:

account for clustering.

Possible techniques depend on design:

  • cluster-robust standard errors;
  • mixed model;
  • GEE;
  • cluster-level analysis.

169. When a specific method is necessary | Say why

If one technique follows from the design or journal standard, explain the reason rather than presenting personal preference as law.

170. Reviewer-as-hidden-coauthor pattern 1 | “Rewrite the Introduction around my theory”

Ask instead:

The Introduction currently presents Theory A as the only explanation, but Theory B predicts the same pattern and is directly relevant to interpretation. Please acknowledge this competing account and clarify whether the current design distinguishes them.

171. Reviewer-as-hidden-coauthor pattern 2 | “Collect my favourite outcome”

Ask whether the missing outcome is necessary to support the manuscript’s current claim.

172. Reviewer-as-hidden-coauthor pattern 3 | “Repeat the whole study in my preferred population”

If current population supports the stated scope, broader replication may be future work.

173. Reviewer-as-hidden-coauthor pattern 4 | “Use my statistical framework”

If current framework is valid, framework preference should not become a major concern.

174. Reviewer-as-hidden-coauthor pattern 5 | “Cite my literature cluster”

Ask whether omission materially distorts the field state.

175. Reviewer-as-hidden-coauthor pattern 6 | “Add a second paper’s worth of mechanism experiments”

Request only what is necessary to support the current mechanism claim—or ask authors to narrow it.

176. Claim reduction is a legitimate repair | Not every weakness needs more data

Reviewer:

The manuscript concludes that the intervention builds long-term competence, but follow-up ends at four weeks. The authors could address this either by restricting the conclusion to four-week performance/persistence or by providing longer-term evidence if available.

177. Give authors more than one valid repair route when possible

This respects authorship and reduces unnecessary work.

178. “Please justify or revise” | Useful reviewer move

Please justify the use of this construct label or revise it to match the validated scale.

179. “Please clarify or narrow” | Useful scope move

Please provide evidence supporting the population-wide claim or narrow the conclusion to the sampled programme.

180. “Please analyse or state as limitation” | Useful uncertainty move

Please assess whether attrition differs by group; if the available data cannot support this analysis, state clearly how differential attrition could affect interpretation.

181. Avoid pseudo-choice when only one repair is scientifically valid

If the primary analysis is simply wrong, be direct.

182. Direct does not mean rude

The current analysis does not account for cluster randomisation and therefore understates uncertainty. This needs correction before the primary effect can be evaluated.

183. Major comment format | Problem → consequence → minimum repair

Example:

Major 1 — Outcome hierarchy. The protocol identifies one-week independent revision as the primary outcome, but the Abstract and Discussion foreground a positive secondary confidence measure while the primary outcome is not clearly reported. This changes the evidence hierarchy. Please restore the primary outcome to the Abstract/Results/Discussion and label confidence as secondary.

184. Minor comment format | Location → issue → fix

Minor 3 — Figure 2 caption: please define the error bars and state the analysed n for each group.

185. Optional comment format | Label it optional

Optional presentation suggestion: Figure 3 may be easier to read if the two time points are shown as separate panels, but this does not affect my scientific assessment.

186. Do not hide a major concern inside “perhaps”

Over-soft:

Perhaps the authors might consider whether clustering could be an issue.

If clustering invalidates uncertainty, say so.

187. Do not inflate a minor concern with catastrophic language

Overstated:

The manuscript is fundamentally flawed because Figure 2 does not define error bars.

188. Calibrate comment force to scientific consequence

189. Explain why the issue matters | Especially for non-obvious statistical concerns

Authors can respond intelligently when they understand the inference at risk.

190. Avoid unexplained jargon as reviewer authority

Weak:

Collider bias.

Better:

The adjusted model includes post-treatment engagement, which may be influenced by the intervention and outcome-related factors. Conditioning on it could induce bias; please justify this adjustment or report an analysis without the post-treatment variable.

191. Give location references | Reduce author search cost

Use:

  • page;
  • line;
  • section;
  • table;
  • figure;
  • equation.

192. Do not line-edit the entire manuscript in the review body

A few representative language examples are enough when the issue is systematic.

193. If terminology drift is systematic | Give one rule and several examples

The manuscript alternates among “confidence,” “self-efficacy” and “competence,” but the instrument appears to measure self-efficacy. Please define the construct once and use a stable term throughout (e.g., Abstract, p. 2; Methods, p. 8; Figure 3; Discussion, p. 17).

194. If numerical inconsistencies are systematic | Ask for global reconciliation

Participant counts differ across Abstract (n=126), Methods (n=120) and Figure 1 (n=112). Please reconcile the recruitment/randomisation/analysis denominators throughout rather than correcting one location only.

195. If causal overstatement is systematic | Ask for global causal-language audit

The design is observational, but causal verbs (“causes,” “leads to,” “improves”) appear in the Title, Abstract and Discussion. Please revise these claims globally to match the design or provide a causal-identification argument.

196. Global comments should name the surfaces affected

Authors need to know this is not a one-sentence fix.

197. Avoid “etc.” in major comments | Be complete enough to act

198. Reviewer comment should not require mind-reading

Weak:

The novelty is unclear.

Better:

The manuscript claims to be the first delayed-transfer study, but Smith et al. and Lee et al. both include delayed outcomes. Please revise the priority claim and specify the narrower contribution of the current unseen-task design.

199. Reviewer should not require impossible perfection | Every paper has residual uncertainty

The question is whether uncertainty is compatible with the claim and journal standard.

200. The “would this change my recommendation?” test

If a requested experiment would not change your recommendation or the manuscript’s validity, reconsider whether it belongs as a requirement.

201. The “could claim narrowing solve this?” test

Before asking for expensive new data, test whether more accurate scope resolves the issue.

202. The “is this already another paper?” test

If the request introduces a new population, mechanism, intervention and outcome, it probably is.

203. The “am I protecting my theory?” test

If you dislike the paper because it contradicts your model, ask whether the evidence is actually invalid.

204. The “am I protecting my citations?” test

Would you still request the citation if it were not yours?

205. The “am I punishing language?” test

Would you make the same scientific recommendation if the prose were polished?

206. The “am I overvaluing novelty?” test

Replication, negative results, validation and methodological correction can be valuable contributions.

207. The “am I overvaluing significance?” test

p<.05 is not a journal-worthiness criterion by itself.

208. Actionable comment for causal overclaim | Model

The cross-sectional design does not establish temporal order, so the statements that X “causes” Y in the Abstract and Discussion are stronger than the design supports. Please revise to association language or provide a justified causal-identification framework.

209. Actionable comment for null/equivalence error | Model

The non-significant difference is interpreted as evidence that the two methods are equivalent. The reported interval remains compatible with meaningful differences in either direction. Please avoid equivalence language unless an equivalence/non-inferiority framework is justified.

210. Actionable comment for subgroup error | Model

The manuscript concludes that the intervention works only in the low-baseline group because the estimate is significant there but not in the high-baseline group. Please test the group-by-baseline interaction directly and present subgroup findings as exploratory if they were not prespecified.

211. Actionable comment for mechanism overclaim | Model

The Discussion attributes the effect to reduced cognitive load, but cognitive load was not measured or manipulated. Please frame this as a possible explanation rather than an established mechanism, or provide direct mechanism evidence if available.

212. Actionable comment for scope inflation | Model

The sample consists of advanced adult volunteers from one high-support programme, while the Conclusion generalises to bilingual learners broadly. Please restrict the empirical claim to the sampled population/context and identify broader transportability as future research.

213. Actionable comment for measure drift | Model

The scale is described as “engagement,” but its items appear to measure attendance and completion behaviour only. Please define the construct more narrowly or justify how the instrument captures broader engagement.

214. Actionable comment for missing-data risk | Model

Twelve-week attrition is 24%, and the manuscript does not report whether missingness differs by group or baseline performance. Please describe the missing-data pattern and assess sensitivity to plausible assumptions, or explain how attrition limits the persistence claim.

215. Actionable comment for selection bias | Model

Programme adoption was voluntary and adopter firms had higher baseline autonomy. This makes the causal productivity claim difficult to support. Please address confounding analytically where justified and revise the causal language to reflect residual selection uncertainty.

216. Actionable comment for data leakage | Model

The feature-selection step appears to use the full dataset before the train/test split, which may leak test-set information into model development. Please confirm the pipeline and, if necessary, repeat feature selection within the training data before evaluating held-out performance.

217. Actionable comment for external validation | Model

The model is described as clinically generalisable, but all evaluation uses an internal split from the development site. Please restrict the generalisability claim or provide external validation in an independent setting if available.

218. Actionable comment for qualitative overfrequency | Model

The Conclusion states that “most students prefer X,” but the purposive qualitative sample was designed to explore experience rather than estimate prevalence. Please report the theme as an observed pattern in this sample rather than a population-frequency claim.

219. Actionable comment for review heterogeneity | Model

The pooled effect is presented as a single general treatment benefit despite substantial heterogeneity across populations and outcome definitions. Please discuss whether the pooled estimate is an appropriate summary and identify the main sources/boundaries of heterogeneity.

220. Actionable comment for novelty overclaim | Model

The manuscript’s contribution appears to be external validation rather than first demonstration of the method. Please revise the novelty framing accordingly and compare directly with the prior development studies.

221. Actionable comment for figure problem | Model

Figure 4 combines two outcomes with different units on one y-axis, making magnitude comparisons difficult to interpret. Please separate the metrics or use clearly defined axes so the visual does not imply a common scale.

222. Actionable comment for protocol mismatch | Model

The registry lists 12-week revision as the primary outcome, whereas the manuscript identifies one-week confidence as primary. Please explain the change, preserve the original outcome hierarchy transparently and update the interpretation accordingly.

223. Actionable comment for preregistration/HARKing | Model

The Introduction presents baseline proficiency moderation as a predicted hypothesis, but the preregistration does not appear to include it. Please clarify whether this hypothesis was prespecified; if it emerged after analysis, label it exploratory and avoid retrospective prediction language.

224. Actionable comment for literature imbalance | Model

The Introduction cites only studies supporting Theory A, although several well-known studies report null or opposing findings. Please represent the evidence state more evenly so the research gap does not depend on selective literature framing.

225. Actionable comment for policy leap | Model

The study demonstrates short-term efficacy in a high-support setting, but the Conclusion recommends system-wide adoption without implementation, cost or long-term evidence. Please separate the empirical finding from the broader policy decision and scale the recommendation accordingly.

226. Actionable comment for harm omission | Model

The intervention improved the target outcome but adverse-event rates were higher in the intervention group. Because this trade-off is decision-relevant, the Abstract and Conclusion should represent both benefit and harm rather than reporting benefit alone.

227. Actionable comment for robustness language | Model

The manuscript calls the result “robust,” but the sensitivity analysis in Table S5 reverses the direction under a plausible specification. Please remove or qualify the robustness claim and discuss dependence on model choice.

228. Actionable comment for feasibility/effectiveness drift | Model

The study is designed as a feasibility pilot, but the Conclusion states that the intervention is effective. Please keep the primary conclusion on recruitment/retention/delivery feasibility and treat outcome trends as exploratory unless the study was powered and designed for effectiveness.

229. Actionable comment for replication contribution | Model

The manuscript is framed as lacking novelty because it replicates a prior study, but the independent replication itself appears to be the central contribution. Please state clearly which aspects are direct replication, which are extensions and how the result updates confidence in the original effect.

230. Constructive rejection comment | Model

The central claim is that X causes Y, but the study measures X and Y at the same time in a convenience sample and lacks a design or analysis that can establish temporal order or address major confounding. Narrowing the language to association would materially change the paper’s stated contribution, and a causal test would require a different study. For that reason, I do not think the current manuscript can support its central claim in revision.

231. Constructive major-revision comment | Model

The primary effect may be interpretable, but the current model ignores instructor clustering and the Conclusion overgeneralises beyond the sampled programme. Both issues appear repairable using the existing data and revised scope language. I therefore see the manuscript as potentially suitable after major revision.

232. Constructive minor-revision comment | Model

The main design and conclusions are supported. My remaining concerns are reporting-level: define the error bars, reconcile one denominator discrepancy and clarify whether the exploratory analysis was prespecified.

233. “Authors should…” can sound unnecessarily commanding | Use when obligation is real

Alternatives:

  • Please clarify…
  • The manuscript would be stronger if…
  • This claim requires…
  • To support X, the authors would need…
  • Please justify or revise…

234. But do not hide requirements behind politeness

If a change is required for validity, say that clearly.

235. “Consider” signals optionality | Use intentionally

The authors may wish to consider…

should not be used for fatal flaws.

236. Avoid sarcasm

Weak:

Surely the authors are aware that correlation is not causation.

Strong:

The cross-sectional association does not establish causal direction; please revise the causal wording accordingly.

237. Avoid rhetorical questions that shame

Weak:

How could the authors possibly conclude this?

238. Convert rhetorical question to diagnostic statement

The conclusion is not supported by the reported analysis because…

239. Avoid mind-reading

Do not write:

The authors ignored this result because it was inconvenient.

Write:

The null primary result is not discussed, while the secondary positive result is emphasised. Please restore the primary outcome to the interpretation.

240. Part III operating rule | A reviewer comment should constrain truth, not authorship

Tell authors what scientific condition must be satisfied. Do not force them to produce the paper you personally would have written unless your requested change is genuinely necessary to make their own paper valid.

Reviewer 最应该控制的是 truth conditions,不是作者的创作控制权。指出必须解决的问题,但不要把“我的研究偏好”伪装成“你的论文必须这样写”。


Part III turns diagnosis into actionable comments. Part IV will assemble those comments into a complete reviewer report: opening summary, strengths, major concerns, minor concerns, confidential editor comments and a recommendation whose severity matches repairability.

Part IV — Assemble the reviewer report and calibrate the recommendation | 第四部分:把 comments 组装成真正可用的 reviewer report

241. A reviewer report should be navigable in five minutes

An editor should quickly see:

  • what the paper does;
  • what is strong;
  • what threatens validity;
  • what is repairable;
  • what recommendation follows.

242. Default report architecture

  1. Short manuscript summary.
  2. Overall assessment / strengths.
  3. Major comments.
  4. Minor comments.
  5. Confidential comments to editor, if needed.
  6. Recommendation.

243. Summary paragraph | Show understanding without rewriting the abstract

Useful structure:

This manuscript examines in [population/context] using [design]. The authors report [main result] and conclude [main claim]. The central contribution appears to be [contribution].

244. Summary should be neutral

Do not begin with verdict language.

Weak:

This flawed manuscript attempts to…

245. Summary is also a fairness check

If the authors would not recognise your summary as their paper, you may be reviewing a straw version.

246. Overall assessment | Separate value from problems

Example:

The question is relevant and the delayed-transfer outcome is a useful contribution. The randomised design is a strength. However, the primary uncertainty estimate does not account for instructor clustering, and the manuscript repeatedly expands the measured revision outcome into broader claims about writing independence. These issues appear repairable with re-analysis and tighter scope.

247. Strengths section | Protect what should remain

Possible strengths:

  • important question;
  • strong design;
  • transparent negative result;
  • external validation;
  • rare sample;
  • good measurement;
  • open data/code;
  • clear reporting;
  • useful replication;
  • careful uncertainty.

248. Do not use praise as a sandwich technique mechanically

The strength section should inform editorial judgement, not merely soften criticism.

249. Major comments should be numbered

Recommended:

Major 1. Primary analysis and clustering

Major 2. Outcome construct and claim scope

Major 3. Twelve-week attrition and persistence

250. Major comments should be ordered by scientific consequence

Do not start with formatting if the causal design is invalid.

251. One major comment may contain linked subpoints

Example:

Major 1a — cluster structure.

Major 1b — updated uncertainty.

Major 1c — downstream Abstract/Conclusion language.

252. Avoid 25 “major” comments

If everything is major, priority disappears. Consolidate comments that arise from the same scientific object.

253. Global concern should be one major comment, not repeated six times

Example:

Causal language exceeds the observational design across Title, Abstract, Discussion and Conclusion.

254. Minor comments should still be precise

Examples:

  • define acronym;
  • reconcile one label;
  • fix caption;
  • state n;
  • correct reference;
  • clarify timing;
  • remove redundant paragraph.

255. Minor comments can be grouped by section

Minor — Abstract

Minor — Methods

Minor — Figures

256. Separate optional presentation suggestions

This prevents authors from treating every preference as mandatory.

257. Confidential editor comments | Use sparingly

Appropriate topics may include:

  • recommendation and repairability;
  • suspected misconduct/integrity issue;
  • conflict concerns;
  • prior knowledge of related submission;
  • journal fit;
  • need for additional specialist review.

258. Confidential comments should not contain hidden scientific attacks

If a methodological weakness justifies rejection, authors normally need to see it.

259. Example confidential note: statistical expertise

I can assess the educational design and interpretation, but I am not confident evaluating the proposed causal mediation model. If the manuscript is otherwise considered further, I recommend additional statistical review of that component.

260. Example confidential note: suspected integrity issue

I noticed apparent duplication between Figure 3 and a previously published image. I have not contacted the authors or investigated independently; I flag this for editorial assessment.

261. Do not accuse authors of fraud in author-facing comments without editorial process

262. Recommendation calibration | Recommendation follows repairability

RecommendationTypical logic
Accept / minorcore validity intact; remaining issues local/reporting
Major revisionimportant repairable scientific issues remain
Rejectcentral claim unsupported and cannot be repaired within current study, or unsuitable for venue

263. “Accept” should not mean “perfect”

Few papers are perfect. The question is whether remaining issues affect validity, interpretation or journal criteria.

264. Minor revision | Core evidence already works

Minor revision should not require re-running a central analysis that could overturn conclusions.

265. Major revision | Central paper may survive, but important uncertainty remains

Examples:

  • cluster correction;
  • missing-data sensitivity;
  • primary outcome hierarchy;
  • construct correction;
  • scope narrowing;
  • new control feasible from existing data;
  • substantial literature reframing.

266. Reject | Do not use rejection as punishment for difficulty

Reject because the paper cannot meet criteria with reasonable revision, not because revisions would be inconvenient to review.

267. Reject for non-repairable design mismatch | Example

The manuscript’s primary contribution is causal, but the available cross-sectional data cannot establish temporal order or address major confounding. Narrowing to association would remove the stated contribution. A fundamentally different study is required.

268. Reject for venue fit | Distinguish quality from fit

The study appears methodologically sound, but its contribution is specialised and may not meet this journal’s stated breadth/priority threshold.

269. Do not tell authors “not impactful enough” without explaining venue criterion

270. Recommendation should not contradict author comments

If author-facing report says “all concerns are minor,” confidential recommendation should not be “reject for serious methodological flaws.”

271. Editor may disagree with your recommendation | That is normal

Your job is to provide reasoning, not control the final decision.

272. Recommendation language | Calibrated examples

Minor: “The central design and conclusions are supported; I recommend minor revision to address reporting and terminology issues.”

Major: “I recommend major revision because the clustering and outcome-scope concerns may materially affect uncertainty and interpretation but appear repairable.”

Reject: “I recommend rejection because the central causal claim cannot be evaluated from the current design and would require new data rather than revision.”

273. Do not include recommendation in author comments if journal keeps it confidential

Follow journal instructions.

274. Reviewer report tone | Firm + neutral

Good report tone is:

  • specific;
  • direct;
  • non-personal;
  • evidence-based;
  • proportionate;
  • repair-oriented.

275. Too harsh | “This analysis is nonsense.”

276. Too vague | “The analysis needs improvement.”

277. Calibrated | “The current analysis treats repeated observations as independent, which understates uncertainty. Please use a model that accounts for within-participant dependence or justify the current approach.”

278. Tone does not require endless hedging

Clear criticism can be respectful.

279. Respect is not agreement

You can recommend rejection in a respectful report.

280. Constructive does not mean “find a way to accept”

A constructive report may explain why the current study cannot support the paper.

281. Avoid praise–criticism–praise formula if artificial

Use scientific structure, not customer-service scripting.

282. First sentence after summary can state overall assessment

The question is worthwhile and the dataset is valuable, but two issues currently prevent evaluation of the primary conclusion.

283. Then list the two issues

This helps editor and authors orient quickly.

284. Reviewer should not overwhelm authors with all possible improvements

Prioritise changes needed for validity and interpretation.

285. “I have 40 suggestions” can be a sign of poor prioritisation

Ask how many are truly necessary.

286. Combine related line edits into a pattern

Instead of 18 comments on terminology, give one global terminology rule with examples.

287. Use examples, not exhaustive copyediting

288. If paper is already publishable | Do not create work to justify your role

A reviewer can say the manuscript is strong with only a few minor comments.

289. The “reviewer value” fallacy | Value is not proportional to criticism volume

A concise report that catches one central issue can be more valuable than a ten-page list.

290. The “complexity” fallacy | Sophisticated request is not automatically a good request

Simple scope correction may solve a problem better than an elaborate new model.

291. The “novelty policing” fallacy | Not every venue requires maximal novelty

Assess against journal criteria and field value.

292. The “significance policing” fallacy | Null findings are not grounds for rejection by themselves

293. The “English policing” fallacy | Do not downgrade sound science because authors are non-native writers

Evaluate clarity and meaning, not accent prestige.

294. The “methodological maximalism” fallacy | More controls are not always better

Extra controls can introduce bias or answer a different question.

295. The “data maximalism” fallacy | More experiments are not always needed

Ask whether current paper’s claim can be supported with proper scope.

296. The “theory maximalism” fallacy | More theories can dilute rather than improve

Require relevant competing accounts, not every possible framework.

297. Re-reviewing a revision | Focus on whether concerns were addressed

COPE advises reviewers, where possible, to review revisions/resubmissions they previously reviewed. The second-round job is not to invent a brand-new list unless new issues arise.

298. Round 2 fairness | Accept scientifically justified author disagreement

If authors did not follow your exact requested method but solved the underlying problem validly, assess the solution fairly.

299. Do not move the goalposts | Revision should converge

Weak reviewer behaviour:

R1: request cluster correction.

Authors correct clustering.

R2: now demand a new population without explaining why it is necessary.

300. New serious issue discovered in revision | Raise it, but explain why it is new

301. Re-review author disagreement | Evaluate reasons, not obedience

Springer Nature guidance for subsequent reviews emphasises focusing on whether prior issues were addressed and evaluating authors’ reasons fairly when they did not follow suggestions.

302. Re-review response letter | Verify every claimed change

Check manuscript, not only rebuttal.

303. Re-review new analyses | Treat as evidence, not reviewer victory

If the new analysis weakens the original result, the revised paper should reflect that.

304. Re-review scope narrowing | Can be a complete solution

Authors do not always need new data if they remove unsupported claims.

305. Re-review declined experiment | Ask whether the paper is still valid within narrowed scope

306. Re-review changed primary conclusion | Tell editor explicitly

The paper may now be publishable for a different, narrower reason.

307. Confidential comment should state recommendation logic, not repeat report

Example:

The clustering and outcome-scope concerns are major but appear repairable with existing data and revised claims; I therefore recommend major revision rather than rejection.

308. Confidential comment can state expertise limit

309. Confidential comment can request specialist review

310. Confidential comment can flag sensitive ethics issue

311. But author-facing science should remain complete

312. Reviewer report final paragraph | Optional

You may end with:

Addressing the major issues above would allow the central contribution to be evaluated more reliably.

313. Avoid threatening closure

Weak:

Until these points are fixed, the manuscript is unacceptable.

314. Avoid bargaining language

Weak:

If the authors cite X and run Y, I will recommend acceptance.

315. You cannot promise acceptance

Editor decides, and revisions may reveal new evidence.

316. The report should stand alone | Another editor should understand it

Do not rely on private notes or memory.

317. Report length should match complexity

Short simple paper may need a short review. Complex multidisciplinary work may require more detail.

318. Long review should still be hierarchical

Summary → major → minor.

319. The editor’s workload matters | Make your judgement legible

Editors may handle many manuscripts. A reviewer who clearly prioritises concerns provides more useful decision support.

320. Part IV operating rule | Recommendation severity must match repairability

Your report should make it obvious why each issue matters, whether it can be repaired, and how that repairability leads to your recommendation. Do not hide a reject recommendation behind minor comments, and do not demand a new research programme under the label “major revision.”

Recommendation 的严重程度必须和 repairability 对得上。不要用 minor comments 支撑 reject,也不要把“重新做一个研究项目”伪装成 major revision。


Part IV assembles the report. Part V will apply the system to a full fictional manuscript and produce a complete peer-review report with summary, strengths, major comments, minor comments, confidential editor note and recommendation.

Part V — Full worked peer review report | 第五部分:完整 Peer Review Report 示例

All study details, authors, data and reviewer comments in this case are fictional teaching material.

321. The manuscript under review | Fictional study

Title:

AI Writing Assistants Build Lasting Independent Writing Ability in University Students

Design:

180 university students are randomised to an AI-assisted drafting tool or a standard word processor for a two-hour writing session.

Primary registered outcome:

unaided essay quality four weeks later.

Secondary outcomes:

  • assisted essay quality during intervention session;
  • completion time;
  • self-reported confidence;
  • unsupported-claim acceptance rate.

322. Fictional Results

Immediate assisted essay quality:

AI group +4.5 rubric points.

Completion time:

AI group 18% faster.

Four-week unaided primary outcome:

difference +0.8 points, 95% CI −1.9 to 3.5.

Confidence:

AI group higher by 6 scale points.

Unsupported claims accepted:

14% of generated factual claims accepted without correction.

323. Manuscript narrative

The Abstract foregrounds immediate quality and speed.

The four-week primary outcome appears late in Results.

The Discussion states that AI “develops independent writing ability by reducing cognitive load.”

The Conclusion recommends widespread university adoption.

324. First reviewer task | Summarise neutrally

This manuscript reports a randomised comparison of AI-assisted drafting and standard word processing among 180 university students. The registered primary outcome is unaided essay quality four weeks after the intervention; secondary outcomes include immediate assisted quality, completion time, confidence and unsupported-claim acceptance. The authors report clear immediate assisted-performance and speed advantages, but the four-week unaided difference is small and imprecise. They conclude that AI develops lasting independent writing ability and recommend broad university adoption.

325. Summary fairness check

The summary does not call the paper good or bad. It makes the evidence hierarchy visible.

326. Strengths

Scientific strengths include:

  • random allocation;
  • prespecified delayed primary outcome;
  • important distinction between assisted and unaided performance;
  • measurement of speed;
  • factual-reliability outcome;
  • practically relevant topic.

327. Overall assessment

The study addresses an important question and the delayed unaided primary outcome is a notable strength. However, the manuscript’s current narrative does not preserve the registered outcome hierarchy: it foregrounds immediate assisted gains while the primary four-week unaided estimate is small and imprecise. The Discussion also attributes the effect to reduced cognitive load without measuring cognitive load, and the final recommendation does not represent the factual-reliability trade-off. I view these as major but largely repairable issues of interpretation and reporting rather than failures of the randomised design itself.

328. Major Comment 1 — Restore the primary outcome hierarchy

The registry identifies unaided four-week essay quality as the primary outcome, yet the Abstract and Discussion foreground the immediate assisted quality gain and describe the intervention as improving independent writing. The primary unaided estimate (+0.8 points, 95% CI −1.9 to 3.5) does not provide clear evidence of a four-week independent-writing benefit. Please restore the registered primary outcome to the Abstract, Results ordering, Discussion opening and Conclusion. Immediate assisted quality and speed can remain important secondary findings but should be labelled as such.

329. Why Major 1 is major

The issue changes the central answer to the research question.

330. Minimum sufficient repair for Major 1

No new experiment is necessary if the authors accurately distinguish:

  • tool-assisted performance;
  • speed;
  • delayed independent performance.

331. Reviewer should not demand a six-month follow-up here

The paper can become valid by narrowing its conclusion to the measured evidence.

332. Major Comment 2 — Construct inflation

The manuscript repeatedly uses “independent writing ability,” but the primary measure is one unaided essay at four weeks. This is a narrower operationalisation than broad writing ability. Please either justify the construct interpretation or revise the language to “four-week unaided essay performance” / “delayed independent writing performance” as appropriate.

333. Why Major 2 is major

The outcome label determines what readers believe has been learned.

334. Major Comment 3 — Unsupported mechanism

The Discussion states that AI improves independent writing “by reducing cognitive load,” but cognitive load was neither measured nor manipulated. The immediate speed gain is compatible with several mechanisms, including reduced drafting effort, language-generation support or direct substitution of model text. Please present reduced cognitive load as a hypothesis rather than an established mechanism, or provide direct evidence if such a measure was collected.

335. Why not demand a cognitive-load experiment?

The current paper does not need to become a mechanism paper if the mechanism claim is narrowed.

336. Major Comment 4 — Benefit–risk trade-off is incomplete

The AI group produced higher-rated assisted essays more quickly, but participants also accepted 14% of generated factual claims that were independently classified as unsupported. This reliability outcome is decision-relevant and should appear in the Abstract and Conclusion. The current recommendation for broad adoption represents the productivity benefit without the factual-risk trade-off.

337. Repair for Major 4

Represent:

  • assisted quality;
  • speed;
  • independent transfer uncertainty;
  • factual reliability.

338. Major Comment 5 — Recommendation exceeds evidence

The Conclusion recommends widespread university adoption, but the study evaluates a single two-hour session and does not establish durable independent-writing benefit, long-term factual reliability, implementation burden or academic-integrity consequences. Please separate the empirical findings from the broader adoption decision. A more proportionate implication would be controlled implementation testing with delayed unaided outcomes and factual-reliability monitoring.

339. Why Major 5 can be repaired without new data

Policy scope can be narrowed.

340. Major Comment 6 — Clarify unsupported-claim measurement

The Methods do not provide enough detail to evaluate the 14% unsupported-claim acceptance outcome. Please define how claims were sampled, what counted as “unsupported,” whether fact-checkers were blinded to condition, how disagreements were resolved and what denominator the 14% represents.

341. Major or minor?

Major because the reliability trade-off is central to the practical interpretation.

342. Major Comment 7 — Analysis of the primary outcome

Please report the full primary model specification, including baseline adjustment, missing-data handling and uncertainty estimation. The manuscript currently reports only a group mean difference, making it difficult to determine whether the registered primary analysis was followed.

343. Do not prescribe a model before seeing the protocol

The review asks for alignment with the registered plan and transparent analysis.

344. Major Comment 8 — Confidence is secondary and should remain secondary

The confidence increase is interesting but should not be used as evidence of improved competence or independent ability. Please preserve the distinction between self-perception and objective performance in the Abstract and Discussion.

345. This is a construct-separation comment, not a demand to remove confidence

346. Minor Comment 1 — Title

The current title (“Builds Lasting Independent Writing Ability”) exceeds the delayed primary evidence. Please revise it to represent assisted performance and/or the registered delayed outcome without implying durable ability.

347. Title could be major | Why placed minor here?

The title is downstream of the major outcome-hierarchy/construct concerns. Fixing those will determine the title. It does not need a separate scientific major issue.

348. Minor Comment 2 — Abstract sample description

Please specify the university-student sample and the four-week timing of the primary outcome in the Abstract.

349. Minor Comment 3 — Figure 2 error bars

Please define the error bars and state analysed n for each group.

350. Minor Comment 4 — Terminology

The manuscript alternates among “AI assistance,” “AI feedback” and “AI tutoring,” although the intervention appears to be a drafting tool. Please use a stable intervention label.

351. Minor Comment 5 — Time language

Please replace generic “post-intervention” labels with “immediate assisted task” and “four-week unaided task” where relevant.

352. Minor Comment 6 — Literature framing

The Introduction would benefit from distinguishing studies of assisted output quality from studies of later unaided transfer; these are not the same evidential question.

353. Minor Comment 7 — Confidence scale

Please give the scale range, validation reference and direction.

354. Minor Comment 8 — Unsupported-claim denominator

Please report numerator/denominator alongside the percentage, especially if participants contributed different numbers of generated claims.

355. Optional suggestion — Figure layout

Optional: consider plotting assisted performance and delayed unaided performance in separate panels to reduce the risk that readers interpret them as the same construct.

356. What this review deliberately does not request

  • a new child/adolescent sample;
  • a six-month follow-up;
  • a neurocognitive mechanism experiment;
  • a different AI model;
  • a Bayesian re-analysis solely by preference;
  • citations to the reviewer’s work;
  • rewriting the Introduction around a different theory.

357. Why restraint improves the review

The existing experiment can answer an important question if its claims are scoped correctly.

358. Confidential comment to editor | Example

The randomised design and registered delayed primary outcome make this a potentially useful paper. My main concerns are evidence hierarchy and over-interpretation rather than an irreparable design flaw. The primary unaided result does not support the manuscript’s current claim of lasting independent-writing improvement, but I believe the paper could become publishable if it is reframed around immediate assisted performance, delayed transfer uncertainty and factual-reliability trade-offs. I therefore recommend major revision. I would be willing to review a revision.

359. Why this confidential note works

It tells the editor:

  • recommendation;
  • repairability;
  • central reason;
  • willingness to re-review.

360. Recommendation | Major revision

Reason:

important repairable interpretation/reporting issues; core randomised design remains useful.

361. Why not reject?

The central valid contribution can survive claim narrowing. No fundamentally new study is required.

362. Why not minor revision?

The Abstract, narrative hierarchy, Discussion and Conclusion all need substantive reinterpretation; analysis transparency may also affect the primary result.

363. Full author-facing review | Model

Summary

This manuscript reports a randomised comparison of an AI-assisted drafting tool and a standard word processor among 180 university students. The registered primary outcome is unaided essay quality four weeks later; secondary outcomes include immediate assisted quality, completion time, confidence and unsupported-claim acceptance. The authors find clear immediate assisted-quality and speed benefits, while the four-week unaided difference is small and imprecise. They conclude that AI develops lasting independent writing ability and recommend broad adoption.

Overall assessment

The question is important, and the delayed unaided outcome is a notable strength. The randomised design is also a strong feature. My main concerns involve the hierarchy and interpretation of outcomes rather than the existence of the experiment itself. In particular, the manuscript foregrounds immediate assisted gains while the registered delayed primary outcome does not clearly support independent improvement; it also overstates mechanism and does not fully represent the factual-reliability trade-off. I believe these issues are substantial but repairable.

Major comments

  1. Primary outcome hierarchy. Restore the four-week unaided primary outcome to the Abstract, Results ordering, Discussion and Conclusion; label immediate quality/speed as secondary.
  2. Construct scope. The measured delayed outcome does not by itself establish broad “independent writing ability.” Narrow or justify the construct language.
  3. Mechanism. Cognitive load was not measured. Present it as a possible mechanism rather than an established explanation.
  4. Benefit–risk trade-off. Include unsupported-claim acceptance in the Abstract/Conclusion and interpret it alongside quality and speed.
  5. Adoption recommendation. Scale the recommendation to the evidence; the study does not yet establish durable benefit or routine-setting safety/implementation.
  6. Unsupported-claim outcome. Add enough methodological detail to evaluate how the 14% estimate was produced.
  7. Primary analysis transparency. Report the registered primary model, missing-data handling and uncertainty estimation.
  8. Confidence. Keep self-reported confidence distinct from objective competence or independent performance.

Minor comments

  1. Revise the title after the primary-claim issue is resolved.
  2. Specify sample and four-week primary timing in the Abstract.
  3. Define Figure 2 error bars and analysed n.
  4. Use a stable intervention label.
  5. Use exact time labels rather than generic “post-intervention.”
  6. Separate assisted-output literature from delayed-transfer literature.
  7. Report confidence-scale range/validation.
  8. Report numerator and denominator for unsupported-claim acceptance.

Addressing these issues would allow the manuscript’s useful randomised comparison to be interpreted more reliably.

364. Full report self-audit | Does every major comment threaten an inference?

Yes.

365. Does every major comment offer a repair path?

Yes.

366. Does the report demand unnecessary new experiments?

No.

367. Does the report distinguish primary from secondary outcomes?

Yes.

368. Does confidential editor note contradict author-facing report?

No.

369. Does recommendation match repairability?

Yes: major revision.

370. Part V operating rule | Review the paper that exists

The strongest peer review often makes the paper smaller, clearer and more honest rather than larger. It protects the valid contribution, removes unsupported expansion and asks for only the evidence needed to evaluate the manuscript’s own research job.

最好的 review 往往不是把 paper 变大,而是让它更小、更准、更诚实:保护真正成立的 contribution,去掉 unsupported expansion,只要求完成这篇 paper 自己必须完成的 evidence job。


Part V provides a complete review report. Part VI will convert the entire lesson into reviewer drills, bilingual comment reconstruction, a 100-point review rubric, ethics checklist, final report gate and an independent manuscript-review benchmark.

Part VI — Reviewer operating system, bilingual drills and final quality gate | 第六部分:把 Peer Review 变成稳定、可迁移的专业能力

371. The reviewer operating system | 一页 reviewer OS

  1. Check expertise. Can I assess the manuscript’s core scientific job?
  2. Check conflict and confidentiality. Can I review fairly and securely?
  3. Identify the research question. What is the paper actually trying to know?
  4. Reconstruct the study identity. Population, design, exposure/intervention, comparator, outcome, time, analysis.
  5. Test validity before style. Does the evidence answer the question?
  6. Separate fatal, major, minor and preference.
  7. Write actionable comments. Observation → consequence → repair target.
  8. Minimise scope creep. Ask only for what the current paper needs.
  9. Calibrate recommendation to repairability.
  10. Keep author-facing science complete.
  11. Use confidential editor comments only for genuinely confidential/editorial issues.
  12. On re-review, judge whether concerns were solved—not whether authors obeyed your exact method.

372. The 10-minute first-pass review | 快速定位 paper identity

TimeTask
2 minread title/abstract and write the claimed research question
2 minidentify design, sample, primary outcome and time point
2 minread primary Results and note the strongest evidence
2 minread Discussion/Conclusion and note the strongest claim
2 mincompare claim with design/evidence for first validity mismatch

373. The 30-minute structured review | Major/minor triage

TimeTask
5 minquestion–design fit
5 minsample/measure validity
5 minanalysis/result integrity
5 minclaim/causality/scope audit
5 minrank concerns fatal/major/minor/preference
5 mindraft one-sentence repair target for each major concern

374. The 90-minute deep review | Full expert workflow

TimeTask
10 minjournal criteria + manuscript identity
15 minMethods/design/measurement
15 minanalysis/statistics/data flow
10 minResults, figures, tables, supplement
10 minDiscussion/Conclusion evidence budget
10 minliterature/originality/reporting/ethics
10 minmajor/minor comment drafting
5 minrecommendation and repairability
5 mintone/bias/confidential-editor audit

375. Reviewer notes should precede reviewer prose | Diagnose first, write later

Keep a private scratch table:

IssueEvidenceInference threatenedSeverityRepairable?
cluster ignored6 instructors, ordinary SEuncertaintymajoryes
broad population claimone specialist sitegeneralisationmajoryes, scope
caption missing nFigure 2clarityminoryes
prefer different colourspersonalnonepreferenceoptional

376. Do not draft angry comments directly into the review system

Write the scientific diagnosis first, then reconstruct it in professional language.

377. Bilingual reviewer problem | 中文内部判断常比最终英文更强

Internal Chinese note:

这个结论完全站不住。

Risky English:

This conclusion is completely wrong.

Professional reconstruction:

The conclusion is stronger than the current design supports because temporal order is not established. Please revise the causal language or provide a design-based causal justification.

378. Chinese “这个方法不对” | Identify what is wrong

Weak:

The method is incorrect.

Better:

The primary model treats repeated observations as independent, which may underestimate uncertainty. Please use an analysis that accounts for within-participant dependence or justify the current specification.

379. Chinese “样本太小” | Name the consequence

Possible reconstructions:

  • The sample yields wide intervals for the primary effect, so magnitude remains uncertain.
  • The subgroup analysis contains too few events for a stable interaction estimate.
  • The model has many parameters relative to the available observations and may be overfit.

380. Chinese “没有创新性” | Replace judgement with contribution analysis

Weak:

The manuscript lacks novelty.

Better:

The claimed contribution as the first delayed-transfer study is not supported because earlier work already includes delayed outcomes. The manuscript may still contribute an independent external validation in this population; please reframe the novelty claim around that specific addition.

381. Chinese “建议补实验” | Ask whether experiment is necessary

Professional reconstruction:

The current mechanism claim requires direct evidence that is not provided by the present design. The authors could either add mechanism evidence if already available or narrow the Discussion to present the mechanism as a hypothesis.

382. Chinese “建议扩大样本” | Avoid ritual sample expansion

Ask:

Which inference would a larger sample repair?

If answer is only “more power,” specify whether current precision is inadequate for the manuscript’s claim.

383. Chinese “结果不显著” | Do not turn into “no effect”

Better review:

The estimate is imprecise and the confidence interval remains compatible with both small benefit and small harm; the manuscript should avoid treating non-significance as evidence of equivalence.

384. Chinese “这个解释太牵强” | Compare alternatives

The proposed mechanism is one plausible explanation, but the same pattern is also compatible with X and Y. Because the current study does not distinguish these accounts, please reduce mechanism certainty and discuss the main alternatives.

385. Chinese “作者忽略了” | Avoid motive accusation

Use:

The manuscript does not discuss the null primary outcome in the Discussion.

Not:

The authors intentionally ignored the null result.

386. Chinese “应该引用” | Citation needs scientific reason

The novelty claim depends on the absence of delayed-transfer evidence, but Study X reports a directly relevant delayed outcome. Please consider this work because it changes the field-state statement.

387. Chinese “语言需要润色” | Be specific about scientific clarity

Several central terms change across sections (“confidence,” “self-efficacy,” “competence”), making it difficult to determine whether the same construct is being discussed. Please define and standardise the technical term.

388. Chinese “结论可以更强” | Evidence, not tone, determines strength

The randomised design supports a stronger causal statement about assignment to the intervention on the measured outcome than the current phrase “associated with.” However, the broader mechanism and generalisation claims should remain cautious.

389. Chinese “结论应该保守一点” | Reduce proposition size, not just add hedges

The study measures one-week revision rather than long-term writing ability; please narrow the construct and time horizon rather than merely adding “may” to the existing broad claim.

390. Chinese “这个 reviewer comment 是 major” | State threatened inference

Major because it affects:

  • primary outcome validity;
  • causal inference;
  • estimate uncertainty;
  • construct validity;
  • evidence hierarchy;
  • generalisation;
  • decision-relevant risk.

391. Bilingual reviewer lexicon | Useful neutral verbs

FunctionUseful English
identifyThe manuscript reports / states / defines…
problemThis does not establish / may not support / leaves unclear…
consequenceThis affects / limits / makes it difficult to interpret…
required repairPlease clarify / account for / reconcile / revise / justify…
optionalThe authors may wish to consider…
scopePlease restrict / distinguish / preserve the boundary…
uncertaintyPlease report the estimate and uncertainty / avoid equivalence language…

392. Drill 1 — Major, minor or preference?

Comment candidate:

The authors use blue instead of green in Figure 2.

393. Model answer 1

Preference unless colour creates accessibility/interpretability problems.

394. Drill 2 — Major, minor or preference?

Cluster-randomised study ignores clusters in primary analysis.

395. Model answer 2

Major: uncertainty and inference may be invalid.

396. Drill 3 — Major, minor or preference?

Figure caption omits error-bar definition.

397. Model answer 3

Minor if underlying analysis is correct; major only if the uncertainty itself is unclear/incorrect.

398. Drill 4 — Major, minor or preference?

Reviewer wants a Bayesian model but current frequentist model is valid and answers the question.

399. Model answer 4

Preference, perhaps optional sensitivity if genuinely informative.

400. Drill 5 — Major, minor or preference?

Primary registered outcome is missing from Abstract and Discussion; positive secondary outcome dominates.

401. Model answer 5

Major evidence-hierarchy concern.

402. Drill 6 — Scope creep

Adult study is internally valid. Reviewer wants a child sample.

403. Model answer 6

Usually future research/generalisation, not mandatory revision, unless the manuscript claims children or the venue requires that population.

404. Drill 7 — Mechanism scope

Treatment effect is clear; mechanism unmeasured.

405. Model answer 7

Ask authors to narrow mechanism claim; do not automatically demand a new mechanistic experiment.

406. Drill 8 — Reviewer bias

Manuscript contradicts your published theory.

407. Model answer 8

Assess design/evidence. If valid, disagreement with your theory is not a defect.

408. Drill 9 — Citation request

Your paper is relevant but not necessary to interpret the manuscript.

409. Model answer 9

Do not request citation for visibility. Suggest only if omission materially distorts the field state.

410. Drill 10 — Language

English is awkward but science is understandable.

411. Model answer 10

Do not make language polish a major scientific concern. Mention clarity only where meaning is obscured.

412. Drill 11 — Recommendation

One analysis issue may change the primary interval but can be fixed with existing data.

413. Model answer 11

Major revision is more coherent than minor or reject, assuming other journal criteria are met.

414. Drill 12 — Recommendation

Central causal claim requires temporal data the study never collected; association-only framing would remove stated contribution.

415. Model answer 12

Likely reject/non-repairable in current study, depending venue and whether an association paper remains meaningful.

416. Drill 13 — Confidential editor note

You suspect image duplication.

417. Model answer 13

Flag confidentially to editor; do not investigate independently or accuse authors in author-facing report.

418. Drill 14 — Re-review

Authors declined your suggested propensity-score method but used a valid regression adjustment that solves the confounding concern.

419. Model answer 14

Judge whether underlying problem is solved. Do not insist on your method.

420. Drill 15 — Re-review

Authors narrowed “long-term learning” to “four-week performance” instead of adding six-month data.

421. Model answer 15

If revised claim now matches evidence, scope narrowing can fully resolve the concern.

422. Drill 16 — Null result

Well-designed study finds no clear primary effect.

423. Model answer 16

Do not reject merely because result is null. Assess design, precision, contribution and reporting.

424. Drill 17 — Replication

Paper replicates prior study closely and finds smaller effect.

425. Model answer 17

Replication can be a valuable contribution; assess execution and how it updates evidence.

426. Drill 18 — AI benchmark

Model outperforms baseline on one internal benchmark; authors claim general superiority.

427. Model answer 18

Major scope concern; request benchmark-specific claim or broader external validation if superiority is central.

428. Drill 19 — Qualitative paper

Reviewer wants statistical power calculation.

429. Model answer 19

Likely inappropriate preference/error. Review sampling logic using qualitative standards.

430. Drill 20 — Reporting guideline

CONSORT flow incomplete but core data exist.

431. Model answer 20

Repairable reporting concern; severity depends on whether omissions prevent evaluation of participant flow/attrition.

432. Seven-day reviewer training cycle | 七天训练

  1. Day 1: classify 40 comments as fatal/major/minor/preference.
  2. Day 2: rewrite vague comments into observation → consequence → repair.
  3. Day 3: review question–design–claim alignment in five papers.
  4. Day 4: practise scope-creep restraint and claim-narrowing alternatives.
  5. Day 5: write confidential editor notes and calibrated recommendations.
  6. Day 6: reconstruct Chinese reviewer notes into neutral English.
  7. Day 7: write a complete unseen review report.

433. Twelve-week C1–C2 peer-review progression

WeeksFocusOutput
1–2ethics, expertise, confidentiality, conflictsreview invitation decisions
3–4validity-first readingevidence maps
5–6major/minor/preference calibrationcomment portfolios
7–8actionable reviewer Englishfull author-facing reports
9–10recommendation, confidential comments, re-revieweditor packages
11–12multi-genre independent reviewsubmission-quality reviewer portfolio

434. The 100-point peer-review rubric | Reviewer mastery

DimensionPointsMastery evidence
Expertise/conflict/confidentiality10review accepted ethically and limitations disclosed
Manuscript identity10question/design/evidence/contribution reconstructed accurately
Validity diagnosis20central inferential threats identified correctly
Severity calibration10fatal/major/minor/preference separated proportionately
Actionability15comments state consequence and minimum repair target
Scope restraint10no unnecessary new-paper requests or theory ownership
Tone/fairness/bias control10direct, non-personal, evidence-based and origin/language neutral
Report architecture5summary, strengths, major, minor and editor note are navigable
Recommendation calibration5severity matches repairability and journal criteria
Re-review discipline5authors’ valid alternative solutions accepted; no moving goalposts

435. What 90–100 looks like | C2-ready reviewer

The review identifies the few issues that genuinely determine whether the manuscript’s claims can be trusted. Major comments are scientifically consequential, specific and repairable where possible. Preferences are clearly optional. The recommendation follows from repairability. The reviewer can disagree strongly without attacking authors, can accept a valid solution they did not personally choose, and does not expand the paper merely to demonstrate expertise.

436. What 75–89 looks like | Strong C1 reviewer

The science is mostly well judged, but one local weakness remains: perhaps too many comments are labelled major, a preferred method is presented too strongly, or confidential/editor-facing reasoning is not clearly separated from author-facing science.

437. What 60–74 looks like | Knowledgeable but over-directive reviewer

The reviewer spots real issues but repeatedly tells authors exactly how to redesign the study, adds scope-expanding requests, focuses too much on language or novelty, or fails to distinguish fatal concerns from preferences.

438. Below 60 | Reviewer as adversary or hidden co-author

The report is driven by theory allegiance, stylistic preference, citation self-interest, prestige bias or a desire to reshape the manuscript into another paper. Return to the manuscript identity and validity ladder before reviewing further.

439. Ethics checklist | Before submitting a review

  1. Was I qualified to review the core scientific job?
  2. Did I disclose conflicts or expertise limits?
  3. Did I keep the manuscript confidential?
  4. Did I avoid unauthorised sharing/co-review?
  5. Did I avoid using unpublished material for my own benefit?
  6. Did I avoid uploading confidential content to unauthorised external tools?
  7. Did I avoid personal/ad hominem criticism?
  8. Did I avoid speculation about author motives?
  9. Did I avoid biased judgement based on origin, institution or language status?
  10. Did I avoid coercive self-citation?
  11. Did I report integrity concerns confidentially to the editor rather than investigate independently?
  12. Does my author-facing scientific report contain all substantive scientific reasons for my recommendation?

440. Scientific review checklist | Before submitting

  1. Can I summarise the paper accurately?
  2. Does design answer question?
  3. Does sample support population claim?
  4. Do measures support constructs?
  5. Does timing support temporal/durability claims?
  6. Does analysis estimate the right quantity?
  7. Do primary outcomes retain priority?
  8. Do numbers and displays reconcile?
  9. Does Discussion stay within evidence?
  10. Does Conclusion preserve causality/scope/uncertainty?
  11. Are limitations consequential?
  12. Is contribution/novelty represented fairly?
  13. Are reporting/ethics/data issues addressed?

441. Comment-quality checklist

  1. Does each major comment name a specific problem?
  2. Does it explain the scientific consequence?
  3. Does it identify a minimum repair target?
  4. Does it avoid prescribing unnecessary technique?
  5. Does it give location/examples where useful?
  6. Does it distinguish required from optional?
  7. Does it avoid mind-reading and personal judgement?

442. Recommendation checklist

  1. Does recommendation follow from the report?
  2. Are major issues repairable with current/feasible evidence?
  3. If reject, is the core issue truly non-repairable or venue-incompatible?
  4. If major, am I asking for a realistic revision rather than a new research programme?
  5. If minor, can remaining issues be fixed without changing the central result?
  6. Have I followed the journal’s recommendation categories?

443. Re-review checklist

  1. Did authors address the underlying concern?
  2. If they used a different method, is it scientifically valid?
  3. Did they preserve planned vs exploratory status?
  4. Did new analysis strengthen or weaken the claim?
  5. Did scope narrowing fully solve the original issue?
  6. Am I moving the goalposts?
  7. Is any new serious issue genuinely new?

444. Final independent benchmark | Review an unseen manuscript

Select an unfamiliar open manuscript or fictional paper. Produce:

  1. reviewer expertise/conflict decision;
  2. one-sentence study identity;
  3. neutral summary;
  4. three strengths;
  5. validity map;
  6. up to five major concerns;
  7. up to eight minor comments;
  8. optional suggestions clearly labelled;
  9. confidential editor note;
  10. recommendation and repairability rationale;
  11. 100-point self-score.

445. Constraint for the benchmark | No scope creep

Before including any requested new experiment, answer:

Which current claim becomes invalid without this experiment?

If you cannot answer, move the idea to optional/future-work territory.

446. Bilingual benchmark | Chinese notes → final English review

Write your raw reviewer notes in natural Simplified Chinese. Then convert each into:

observation → consequence → repair target → severity.

Compare the English report with the Chinese notes. Remove:

  • extra certainty;
  • personal judgement;
  • “obviously / clearly wrong” language;
  • commands that are merely preference;
  • unnecessary politeness padding;
  • scope-expanding requests.

447. Final assignment | Complete peer-review package

Produce:

  1. review invitation decision;
  2. conflict/expertise statement;
  3. evidence map;
  4. full author-facing report;
  5. confidential editor note;
  6. recommendation;
  7. self-audit;
  8. second-round review after a fictional author response.

448. Research and reference floor | 研究与参考基础

449. Canonical eduKate route | 研究写作与 peer-review 路线

450. SEO language map | 本课自然覆盖的搜索意图

This lesson naturally serves readers searching for how to peer review a paper, peer review report example, reviewer comments example, major comments minor comments, how to review research paper, reviewer report structure, confidential comments to editor, peer review ethics, reviewer conflict of interest, peer review confidentiality, constructive peer review, how to recommend major revision, how to recommend rejection, actionable reviewer comments, avoid reviewer bias, peer review for non native English speakers, academic English reviewer report, C1 academic writing, C2 academic English, 审稿意见怎么写, peer review 怎么写, 审稿报告, major comments, minor comments, reviewer comments, 审稿伦理, 审稿人意见 and 中文母语学术英语.

451. Final 25-question reviewer gate | 提交 review 前

  1. Am I sufficiently expert for the core scientific job?
  2. Have I disclosed conflicts/expertise limits?
  3. Have I protected confidentiality?
  4. Can I summarise the manuscript neutrally?
  5. Does the design answer the stated question?
  6. Does the sample support the population claim?
  7. Do the measures support the constructs?
  8. Does the analysis estimate the relevant quantity?
  9. Are primary outcomes represented fairly?
  10. Do numbers/tables/figures reconcile?
  11. Does Discussion stay within the evidence?
  12. Does Conclusion stay within causality/time/population bounds?
  13. Have I separated fatal/major/minor/preference?
  14. Does each major comment name a threatened inference?
  15. Does each major comment offer the minimum sufficient repair target?
  16. Am I demanding any experiment that belongs to another paper?
  17. Am I forcing my preferred theory/method/software?
  18. Am I suggesting citations only for genuine scientific reasons?
  19. Have I avoided author-motive speculation and personal criticism?
  20. Have I separated optional suggestions from required changes?
  21. Does my confidential editor note avoid hidden scientific criticism?
  22. Does my recommendation match repairability?
  23. Would I make the same judgement if author/institution/country/language were different?
  24. If this is a re-review, have I avoided moving the goalposts?
  25. Will authors and editor understand exactly what matters most?

452. Final principle | The reviewer protects inference, not ownership

Peer review is at its best when it protects the truth conditions of a manuscript while preserving the authors’ ownership of the research question and intellectual architecture.

A reviewer should be strict about:

  • validity;
  • evidence;
  • uncertainty;
  • scope;
  • transparency;
  • ethics.

A reviewer should be humble about:

  • personal theoretical preference;
  • preferred technique when alternatives are valid;
  • style;
  • citation visibility;
  • desire for a larger paper.

Reviewer 的权力应该用来保护 inference,而不是夺走 authorship。严格地守住 evidence,克制地使用 preference——这才是高质量 peer review。

453. Exit standard | You are ready to move on when…

You can review an unfamiliar manuscript and:

  • identify its research identity before criticising;
  • detect validity-threatening flaws;
  • separate major/minor/preference;
  • write comments with consequence and repair targets;
  • avoid unnecessary new-paper requests;
  • control bias/conflict/confidentiality;
  • calibrate recommendation to repairability;
  • re-review fairly when authors choose a different valid solution.

At that point, you are not merely “giving comments.” You are performing disciplined scientific quality control.


Next lesson reserved | 下一课

EDKS-ADV-ZH-0025 · Lesson No.025 · Give a Research Presentation and Defend It Under Questions Without Overclaiming | 做研究汇报并在问答中守住证据边界

The next lesson moves from written research to live research English: designing a research presentation, compressing methods and results into spoken form, answering hostile or difficult questions, distinguishing “I do not know” from “the data do not establish this,” defending justified choices without becoming defensive, and preserving claim strength under time pressure.

Back to Advanced English Chinese Edition Hub · 返回高级英语中文版主页