How to conduct a systematic literature review is a practical question for university students who want to understand what a body of research really establishes, instead of selecting a handful of papers that support a preferred conclusion. A systematic review uses a defined question, reproducible searches, explicit eligibility rules, documented screening and a careful synthesis of evidence. It can reveal agreement, disagreement, missing research or methodological weaknesses—but only when the reviewer’s method is as transparent as the result.
For graduates researching PRISMA 2020 checklist, systematic review methodology, literature search strategies, screening studies and research evidence synthesis, the main official reference is the PRISMA 2020 statement, a reporting guideline with a 27-item checklist and flow diagrams. NTU Libraries’ systematic review guidance explains how to report information sources, searches and study selection so another reader can examine the process. Importantly, PRISMA primarily improves reporting; it is not, by itself, a complete substitute for designing a valid systematic review.
This guide explains how to formulate a reviewable question, plan a protocol, search databases, remove duplicates, screen eligible evidence, evaluate bias, synthesise findings and report limitations. The numerical examples are explicitly fictional teaching exercises, not real systematic reviews. Students should follow their faculty’s subject-specific standards and seek appropriate library or methodological support where needed.
A systematic review is not a long list of summaries
A conventional essay may select readings to support an argument. A systematic review asks a defined question and attempts to identify relevant studies through a method that can be scrutinised or repeated. Inclusion and exclusion rules should be established before the reviewer knows which results will look favourable.
The objective is not to accumulate as many citations as possible. It is to assemble a defensible body of evidence and explain its strengths and weaknesses. A review that silently ignores contradictory studies cannot claim systematic rigour merely because it includes a flow diagram.
First decide which type of evidence synthesis fits
Systematic reviews often address focused effectiveness, association, diagnostic, qualitative or other questions depending on discipline. Scoping reviews may instead map how a field has been studied and where research is concentrated, without necessarily estimating one intervention’s effect.
These are not interchangeable labels. Read the methods standards and journal requirements relevant to the particular review type. The general PRISMA 2020 statement mainly targets systematic reviews of intervention effects, while extensions and other reporting guidance address different synthesis designs.
A good question determines the whole method
An excessively broad question such as ‘Does technology improve education?’ spans many interventions, subjects, ages and outcomes. A better research question identifies who or what is studied, which intervention or exposure matters, what comparison is relevant and which outcome is important.
Defining scope is not artificially narrowing research to get a desired answer. It creates clear criteria for what evidence can answer the question. A graduate can revise the question during protocol development, but should document material changes once formal screening begins.
PICO is a useful framework for some questions
For a review of an intervention, PICO commonly stands for Population, Intervention, Comparator and Outcome. It can help identify appropriate search terms and eligibility criteria. Not every research question needs a rigid comparator, especially in qualitative or descriptive syntheses.
A student should select a framework suited to the discipline rather than force an unsuitable social or theoretical topic into clinical-trial vocabulary. Ask which elements genuinely determine whether a study can answer the question.
A fictional PICO example
Suppose a fictional educational review asks whether a structured feedback method helps older secondary students revise scientific explanations compared with ordinary comments. Population, intervention, comparator and defined outcome can be specified before searching.
The example does not establish that the feedback method works. The review must still seek studies with appropriate designs and examine their methods. A clear question prevents the reviewer from switching outcomes to whatever one paper happens to report positively.
A protocol records the decisions beforehand
A systematic review protocol can define the question, eligibility criteria, databases, search approach, screening procedure, data fields, risk-of-bias assessment and planned synthesis. Registration may be appropriate in a relevant registry depending on the discipline and eligibility.
The protocol makes later departures visible. A student should record why a search was revised or eligibility narrowed after evidence appeared. A protocol can improve transparency, but simply publishing a protocol does not guarantee the methods will be executed correctly.
PRISMA 2020 is principally a reporting framework
The PRISMA 2020 checklist describes what a completed review should report, including methods, included-study characteristics, risk-of-bias assessments, synthesis, limitations and information about the protocol.
Completing its checklist is not automatic proof of a high-quality review. A team must still use scientifically appropriate methods, independent judgement and reliable sources. Report transparency and methodological quality are related but distinct.
PRISMA-S goes deeper into searching
The PRISMA-S extension contains specific recommendations for reporting systematic literature searches, including details that allow another reader to understand which databases, search expressions and limits were used.
This is especially useful when an academic paper’s final search yields hundreds or thousands of records. Saying ‘we searched Google and PubMed’ usually does not provide a reproducible strategy. The precise database interface, dates and search terms matter.
Search broadly enough to answer the question
Different scholarly databases index different journals and disciplines. A review of educational interventions may require sources beyond those used for biomedical trials, while engineering or computing studies can include relevant proceedings.
Select databases based on the question and available institutional access rather than habit. A librarian can help identify coverage, controlled vocabulary and field-specific sources. The aim is justified coverage, not searching every possible website without a plan.
Record complete database searches
For each source, record the database name, platform where relevant, date searched, exact search expression, limits and number of retrieved records. Search syntax differs by platform, so an expression copied unchanged into another database may not have the same meaning.
Keep the actual strategies in a permitted file or appendix so another researcher can inspect them. A final thesis should not claim an exhaustive search when important sources were omitted without explanation. Document what was really done rather than reconstruct a cleaner history afterward.
Combine concepts with Boolean logic
The terms AND, OR and NOT can combine search concepts, although the exact syntax varies by database. OR often joins synonyms, while AND combines distinct parts of the research question. Quotation marks, wildcards and subject headings are platform-specific.
A student should test whether the search finds known relevant studies and adjust carefully. A query that returns only a few favourable papers may be too restrictive. Conversely, thousands of irrelevant results can indicate that terms lack useful precision.
An illustrative Boolean search
For a fictional question about feedback and student writing, a preliminary search might combine (feedback OR comments) AND (writing OR revision) AND (secondary school OR adolescent). The expression should be translated appropriately for the real database and its controlled vocabulary.
This is a teaching example, not a validated exhaustive search strategy. It may miss relevant synonyms, regional terms or studies describing different educational stages. Use a librarian or appropriate information specialist to refine and document a full search.
Date and language restrictions need justification
A review may have a defensible reason to focus on a recent technology, a period after a policy change or studies available in specified languages. Every restriction can nevertheless exclude evidence and introduce bias.
Record the justification and recognise which populations, publications or findings may have been missed. An arbitrary ‘last five years only’ filter does not automatically make a study scientifically better. The correct time window follows the question.
Grey literature can help reduce bias
Relevant evidence may appear in dissertations, conference records, government reports, trial registries or other material outside standard scholarly journals. Whether to include these depends on the review’s question and protocol.
A published-journal-only search may miss negative, incomplete or less newsworthy outcomes. But grey literature also requires careful source and quality evaluation. A document being publicly downloadable does not prove its methods are reliable.
Keep all retrieved records in a traceable system
Search results can be imported into an authorised reference manager or screening tool. Each record should retain enough identifying information to locate the source and establish where it was discovered.
Do not rely on copying article titles into an unlabelled list. A durable record makes it possible to identify duplicates, explain exclusions and update the search later. Use approved research systems for any sensitive or restricted information.
Remove duplicates before screening
The same article can appear in several databases, sometimes under slightly different records. Deduplication prevents counting it as several independent studies. Keep information about the number of duplicate records removed.
A duplicate record differs from a distinct report of the same underlying study. Two publications from one trial should not automatically be treated as two independent experiments in meta-analysis. Screening needs to track study identity as well as document identity.
Two-stage screening improves efficiency
Reviewers often screen titles and abstracts first, then assess potentially relevant full texts against the same eligibility criteria. Decisions should be documented so a reader can understand why records were excluded.
Do not exclude a study because its result seems unfavourable. Where abstracts are ambiguous, inspect full text rather than guessing. The procedure must follow the protocol and appropriately handle discrepancies, ideally using independent review where possible.
Independent screeners can reduce errors
Many systematic-review methods use two reviewers or an appropriate independent validation process to assess eligibility. This helps detect inconsistent judgements and resolve ambiguity about criteria.
Student resources may be limited, but limited reviewer capacity should be reported rather than concealed. If a single reviewer made all decisions, describe the limitation and any independent checks. Do not invent a second screener or fabricate agreement statistics.
A PRISMA flow diagram explains record counts
The PRISMA 2020 flow diagram tracks the flow from identification through deduplication, screening, retrieval, eligibility assessment and inclusion, with reasons for relevant exclusions.
It is a report of actual record handling, not a decorative graphic. Each count should reconcile with the underlying screening log. Reviewers cannot simply select round numbers that fit the boxes more neatly than the real search.
A fictional screening count example
Imagine a teaching exercise retrieving 180 database records and 20 from other sources. After removing 35 duplicates, 165 distinct records remain for initial screening. If 120 are excluded at title or abstract stage, 45 remain to be sought as full reports.
These figures are invented only to demonstrate record accounting. The team must track reports that could not be retrieved and document full-text eligibility decisions. A real PRISMA diagram should follow the actual template appropriate to database and other-source searches.
Eligibility criteria should be operational
A criterion such as ‘good quality studies’ is too vague without a defined assessment framework. Specify population, design, setting, outcome, publication type and other relevant conditions in ways independent reviewers can apply consistently.
A student should distinguish the decision about whether a study is relevant from the later judgement about its risk of bias. Relevant but weak studies may need to be included and critically assessed rather than silently discarded after the result is known.
Document full-text exclusion reasons
When a report appears promising but fails the inclusion criteria on detailed reading, record the specific reason. Examples include wrong population, unsuitable comparator, no relevant outcome or an excluded study design under the protocol.
Do not mark every inconvenient study as ‘not relevant’ without an explanation. A transparent exclusion log helps another reviewer challenge or reproduce the choice. The PRISMA checklist also asks authors to address studies that might appear eligible but were excluded.
Extract data using defined fields
A data extraction form may include author, publication year, country, participants, design, intervention or exposure, comparator, outcomes, measures, follow-up and limitations. Qualitative reviews can use different appropriate fields.
Pilot the form on a small number of studies before full extraction. Record what a paper actually reports, including missing details, rather than infer favourable data. A consistent codebook prevents reviewers from using the same variable label for different meanings.
An evidence table is more than citation storage
A useful table makes studies comparable by question, design, sample, outcome and main limitations. It should support a reasoned synthesis, not merely list paper titles and authors.
Two studies with identical conclusions may still have very different reliability because of recruitment, measurement or analytical methods. A strong review makes those differences visible. The reader should understand why apparent agreement is or is not persuasive.
Risk of bias asks whether the study’s design could mislead
A study can measure an outcome accurately yet still produce a misleading conclusion because groups differed before an intervention, participants were selected unusually or important outcomes were unreported. Risk-of-bias assessment examines such methodological threats.
Use an appropriate established tool for the study design and question. A randomised trial, observational study and qualitative interview need different evaluative approaches. Do not invent a universal quality score that treats every research method identically.
Risk of bias differs from how prestigious a journal is
An article published in a widely read journal still deserves scrutiny of its methods, while a less prominent study may contain careful evidence. Journal reputation alone is not a legitimate substitute for an explicit risk-of-bias assessment.
A reviewer should document why a particular feature raises concern and how it affects the interpretation. Avoid labelling evidence high-quality merely because the paper is old, widely cited or authored by a familiar academic.
Distinguish reporting quality from study quality
A study that omits important method details can be difficult to assess, but poor reporting is not always proof that the underlying experiment was conducted badly. Conversely, a well-formatted paper can still use biased methods.
Record uncertainty as uncertainty. Where permitted, contacting the original study authors for clarification may help. A systematic review should not silently assume the most favourable interpretation when evidence about a method is missing.
Meta-analysis is a method, not a requirement
A meta-analysis statistically combines suitable quantitative study results. A systematic review may use it when outcomes, designs and populations are sufficiently comparable and the necessary numerical information exists.
If studies are too different, a structured narrative or another appropriate synthesis may be more defensible. Never pool incomparable results simply because a forest plot looks rigorous. A careful explanation of why pooling was inappropriate can be a substantive methodological decision.
The effect measure matters before pooling
For binary outcomes, reviewers may consider measures such as risk ratios or odds ratios; continuous outcomes may require mean differences or standardised effects. Selection depends on the question, design and data.
Different effect measures are not freely interchangeable. The reviewer should identify what one unit or ratio means and how precision was calculated. An apparently impressive percentage difference may represent an uncertain estimate from a small or unrepresentative sample.
Precision is more informative than a point estimate alone
A combined effect estimate describes one summary of the eligible evidence, but its confidence or credible interval shows uncertainty under the statistical model. A narrow estimate and a wide one should not be described with identical certainty.
A study may point toward a benefit while remaining too imprecise to distinguish moderate improvement from little effect. Report uncertainty faithfully. The PRISMA 2020 checklist asks for effect estimates and their precision when applicable.
Heterogeneity can change the meaning of a pooled result
Studies may involve different populations, settings, interventions or measurements. Statistical heterogeneity describes variation beyond what would be expected from sampling error alone under a model, but it cannot explain all practical differences.
A pooled number may conceal that an intervention appears useful in one setting and ineffective in another. Reviewers should inspect potential reasons, the studies’ methods and the limits of subgroup analyses rather than treating one summary as universal truth.
A fictional heterogeneous-evidence example
Imagine two completely invented studies of classroom feedback. One involves older students over a full term, while another uses a single online practice session with younger learners. Both report positive results, but their designs and time horizons differ.
It may be inappropriate to describe them as direct replications. A synthesis could explain what each actually establishes and where further comparable research is required. The example is methodological teaching, not a real claim about an educational intervention.
Narrative synthesis still requires discipline
A narrative synthesis can organise findings by study design, population, outcome and context, then compare patterns and uncertainty. It should not become an unstructured collection of summaries chosen in a persuasive order.
Explain the method used to group studies and how conflicting results were handled. A systematic narrative synthesis is rigorous when its reasoning is transparent and evidence-led, even without a numerical meta-analysis.
Publication bias means missing evidence may matter
Studies with strong or statistically significant outcomes can be more likely to appear in accessible publications than studies showing no clear result. This can distort what a review concludes if its search strategy excludes other relevant reports.
Consider trial registries, dissertations or other eligible sources where appropriate, and discuss whether missing evidence threatens certainty. A claim that all relevant studies have been found is difficult to support without comprehensive searching and honest limitations.
Certainty of evidence is another judgement
Risk of bias evaluates potential problems within studies; certainty of a body of evidence asks how confidently the combined findings answer a particular question, considering other factors such as inconsistency, imprecision and relevance.
Use an appropriate framework for the discipline and review type. A large number of included studies does not automatically make certainty high. Several similarly biased studies may provide less dependable information than a smaller set of sound studies.
Sensitivity analyses test assumptions
A quantitative review may examine whether results change when excluding studies with certain methodological concerns or using different justified analytical assumptions. These checks can reveal whether conclusions depend on one fragile choice.
Report planned analyses and any post-hoc exploration honestly. Do not search through many alternative calculations only to publish the one yielding the most attractive result. A sensitivity analysis should improve understanding, not manufacture significance.
Use a transparent review spreadsheet
Record source, identifier, screening decision, full-text exclusion reason, extracted fields, reviewer checks and relevant methodological concerns in a documented workspace. Keep original study records and processed summaries distinguishable.
A research team may use approved review software, but the software does not make decisions scientifically valid. The logic behind criteria, extraction and synthesis must be explicit. Protect any confidential or licensed material under institutional policies.
What to do with conflicting findings
Contradiction is not necessarily evidence of a flawed review. Researchers may have studied different groups or contexts, chosen different outcome measures or used designs with unequal vulnerability to bias.
Examine whether the studies truly address the same question and how the methods affect comparability. A review that explains disagreement carefully can be more useful than one that forces a single simple conclusion.
The PRISMA diagram needs internally consistent arithmetic
Imagine a fictional search identifies 220 records, removes 35 duplicates and screens 185. If 140 are excluded during title and abstract screening, 45 reports are sought. If five cannot be retrieved, forty reports are assessed for eligibility.
Suppose 25 of those fail criteria with documented reasons, leaving fifteen reports included. The counts should reconcile with the real tracking log. These invented values show bookkeeping only; they do not prove fifteen independent studies were included or that any effect was found.
Reports and studies are not necessarily the same
One trial can produce a registry record, conference abstract and journal article. PRISMA 2020 distinguishes reports from underlying studies because multiple documents may describe the same investigation.
A reviewer should link related reports instead of treating every record as an independent experiment. Otherwise a meta-analysis could accidentally double count participants and make the evidence look more precise than it is.
Update the search before finalising where appropriate
A review prepared over several months may use searches that are old by publication time. Depending on the discipline, journal and protocol, it can be appropriate to update searches before final synthesis.
Record the date and method used for each update and account for newly discovered studies. Do not quietly incorporate a few convenient late papers without applying the same eligibility and quality procedures as the initial search.
Document protocol departures honestly
Research practice sometimes reveals that an outcome is unavailable or that an eligibility criterion was unclear. A necessary amendment should be described, including why it occurred and whether it could affect the conclusion.
The PRISMA 2020 checklist asks authors to report protocol registration and amendments. Transparent change control is stronger than rewriting the original plan as if every decision had been made before the evidence appeared.
Avoid automatic assumptions about registration
Relevant registries can help make a review protocol discoverable, but their scope and eligibility differ by topic and type of review. A student should not assume that every educational or engineering review can be entered into a healthcare-focused registry.
Check registry requirements and any departmental expectations, including whether the review type is eligible. If the protocol is not registered, report that accurately where required. Never invent an identifier to make a review appear more rigorous.
Human-subject ethics and evidence synthesis are separate
A systematic review using only lawfully available published papers may not require the same human-participant ethics process as a new interview study. However, reviews using protected individual participant data or confidential records can raise separate permissions and ethics questions.
Ask the institution when uncertain, particularly for patient-level datasets. Read our research ethics guide for the distinction between review methods and permission to collect identifiable information.
AI tools can assist only within verified rules
Tools may help organise permissible citations or draft search expressions, but generated references and purported screening results can be inaccurate. A review author remains responsible for checking original papers and decisions.
Do not let an AI system invent included studies, PRISMA counts, data extraction or effect sizes. Verify every source and follow the university’s current rules on AI assistance and confidential documents. Automated convenience is not a substitute for independent screening judgement.
How to write a defensible discussion
Begin by stating what the body of included evidence indicates under the chosen methods. Then explain how risk of bias, inconsistency, imprecision and scope limit the conclusion, and identify reasonable implications for further research.
Avoid grand statements that the review settles a complex field forever. A bounded conclusion might say evidence is promising in specified settings but too limited to support widespread application. That level of honesty makes the review more valuable to readers who must make decisions.
A collaborative review needs a clear contribution record
Team members may help design the search, screen records, extract evidence, assess bias, analyse data or draft interpretations. Define responsibilities and resolve disputes through an appropriate process before submitting the review.
A contributor who performed a meaningful task should be acknowledged or considered for authorship under the relevant journal’s rules. Do not promise authorship solely for seniority or remove a colleague because their results were inconvenient. See our research publishing guide.
A six-stage student review plan
First develop a focused question and protocol. Second select sources and build reproducible searches. Third identify and deduplicate records. Fourth apply documented screening and eligibility. Fifth extract data, assess bias and synthesise findings. Sixth report the methods and evidence with appropriate PRISMA guidance.
This sequence is an educational overview, not a prescribed schedule for every review. The volume of literature, methods and team capacity can make a real review a substantial project. Plan academic support and deadlines honestly rather than assume a systematic review can be completed by searching one afternoon.
Frequently asked questions
Is PRISMA a search database? No; it is primarily reporting guidance. Must every systematic review have a meta-analysis? No. Can a student use only five favourite papers? Not while claiming a comprehensive systematic approach without an appropriate search and eligibility rationale.
Does a PRISMA diagram prove research quality? No. Can all review protocols be registered in the same registry? No. Must studies with null results be included? When they meet the stated criteria, results should not be excluded merely because they are unfavourable.
Official resources and connected eduKate reading
Primary sources: PRISMA 2020, PRISMA checklist, flow diagram templates, PRISMA-S and NTU Libraries Systematic Reviews.
Related eduKate guides: Academic Research Writing, Research Data Management and Publishing Research.
Final thought: transparency makes synthesis valuable
A systematic review is useful when readers can see how its question was defined, where evidence was sought, why studies were included, what weaknesses remain and how conclusions follow from the actual findings.
Use PRISMA to improve reporting, but invest equally in appropriate methods and honest decisions. A carefully bounded conclusion can help future researchers more than a confident claim unsupported by the full body of evidence. Reviewed 8 October 2026; reporting guidance and subject-specific methods can evolve.
