社交

A Better Search Strategy for Stronger Systematic Reviews

A Systematic Review Is Won or Lost Before the First Search


A systematic review is a method for defining a question in advance, finding all eligible research, judging how trustworthy each study is, and combining the evidence. It can look rigorous at the end and still have been compromised at the beginning. It may contain a polished flow diagram, hundreds of screened records, tables assessing each study's risk of systematic error, and a carefully calculated meta-analysis that statistically combines results from multiple studies. None of those later steps can recover a relevant study that the search never found. If the missing studies differ systematically from the studies retrieved, the review's conclusion can be confidently precise and still be wrong.


This is why literature searching is not a clerical prelude to the real research. It is one of the review's core methods. A search determines which part of the evidence landscape is allowed to become visible. Its design therefore belongs in the protocol, before the team sees a convenient pattern in the results. A protocol is the advance record of the question, eligibility criteria, sources, search approach, screening process, and analysis plan. It does not freeze a review against every sensible amendment. It makes changes visible, dated, and explainable rather than allowing the method to drift silently toward an attractive conclusion.


A Better Search Strategy for Stronger Systematic Reviews


Major methods guides converge on this principle. The Cochrane Handbook expects searching to be planned during protocol development and recommends information-specialist involvement when the team lacks search expertise(1). The JBI Manual for Evidence Synthesis treats search development as an iterative process of exploration, vocabulary discovery, testing, and translation across sources(2). PRISMA-P asks authors to make the planned information sources and draft strategy visible before the review is complete(10). The practical lesson is simple: do not begin by asking which words to type. Begin by deciding what evidence would answer the question and how you will know it when you see it.


Planning also has to cover the operational trail around a query. The National Academies standards call for a question-specific comprehensive search, information-specialist planning, independent peer review, and line-by-line strategies with dates and dispositions for identified reports(4). CRD guidance likewise treats databases as one component of identification alongside updates, reference management, and document retrieval(5). Both were written for health-care reviews, and CRD's named services may now be dated, but their durable point is that a reproducible search preserves its operating record, not merely its final Boolean string.


Turn the question into a search design


Suppose a research team wants to review whether flexible work arrangements affect employee mental health. That sentence names an interest, not yet a reproducible question. Does “employee” include freelancers and self-employed workers? Does flexible work mean control over start and finish times, remote work, compressed weeks, or any departure from a fixed office schedule? Does mental health mean diagnosed depression, a validated stress scale, burnout, self-reported well-being, sickness absence, or all of these? Are cross-sectional surveys eligible, or only studies that can establish time order? Each answer changes both the evidence that counts and the vocabulary required to find it.


Question frameworks can expose these decisions. PICO, for example, separates a population, an intervention or exposure, a comparison, and an outcome. Other review types may use different frameworks because a qualitative experience, a diagnostic test, or a policy question is not well represented by an intervention template. Even when PICO fits the question, it should not be mistaken for a rule that every element must appear in the database query. A systematic review of PICO's effect on search quality found only three sufficiently relevant comparative studies and could not support a firm practice conclusion(17). Outcome terms are particularly risky: authors describe outcomes in many ways, and some outcomes are absent from titles and abstracts. Requiring one narrow outcome vocabulary can make a query look specific by quietly discarding eligible studies.


The eligibility criteria and the search strategy have related but different jobs. Eligibility criteria define the final boundary of the review. A search strategy is an imperfect detection instrument designed to bring possible members of that set into view. It will usually need to be broader than the final criteria. This distinction explains the central tradeoff between recall and precision. Recall, also called sensitivity in this context, is the proportion of all relevant records that the search retrieves. Precision is the proportion of retrieved records that turn out to be relevant. Increasing recall often produces more irrelevant material to screen. Increasing apparent precision by adding many restrictive concepts can conceal a serious loss of relevant studies.


A Better Search Strategy for Stronger Systematic Reviews


Systematic reviews generally give recall priority because screening can remove an irrelevant record, whereas no later step can screen a record that was never found. But “retrieve everything” is not an operational method. A search must be tested against the question, known relevant studies, likely terminology, and the team's justified resource constraints. The right question is not whether a query returns many records. It is whether the query retrieves the kinds of studies the protocol says matter without imposing restrictions that the evidence cannot survive.


For each essential concept, searchers normally combine free-text language with controlled vocabulary. Free text means the words that authors use in titles and abstracts: flexible schedule, flexitime, remote work, telework, work from home, and so on. Controlled vocabulary is a database's standardized subject-label system, analogous to a carefully maintained library classification. PubMed's Medical Subject Headings, or MeSH, are a familiar example. Controlled terms collect papers that use different surface wording, but they are not enough on their own. New records may not yet be indexed, new concepts may lack a settled heading, and indexers can make different judgments. Free text is also insufficient because spelling, phrasing, abbreviations, and disciplinary vocabulary vary. A defensible strategy uses both.


Synonyms within one concept are usually joined with OR, while distinct concepts are joined with AND. This Boolean structure is easy to state and easy to damage. Missing parentheses can change the logic of an entire search. Quotation marks can force words to appear as a phrase. Proximity operators require terms to appear within a specified distance, truncation symbols retrieve several word endings from one stem, and field codes restrict a term to places such as the title or abstract. All of these work differently across platforms. NOT is especially hazardous because a record containing the excluded term may still report eligible evidence. The objective is not to produce the longest possible string. It is to express the protocol's concepts with enough linguistic coverage and correct database-specific syntax.


Search vocabulary should be discovered, not guessed once. JBI's three-step approach begins with a limited preliminary search, mines relevant titles, abstracts, and index terms, then expands and adapts the strategy across selected databases. A set of known relevant “seed” papers is useful here. If the draft query cannot retrieve them, the team has a diagnostic problem to solve: perhaps an essential synonym is missing, a subject heading was not configured to include its narrower subtopics, a phrase is over-restricted, or the chosen database does not cover that literature. Seed-paper testing does not prove that every relevant paper will be found, but failure to retrieve a central known paper is clear evidence that the strategy is not ready.


Design a discovery system, not a single query


No database is a neutral window onto all scholarship. Databases differ in journal coverage, document types, regional representation, language, indexing depth, update speed, and search functions. Source selection should therefore follow the question and the likely publication routes of its evidence. A health intervention, an education policy, a software-engineering technique, and a humanities interpretation do not leave the same documentary trail.


Empirical comparisons show why convenient one-source rules are unsafe. In a prospective study of 58 biomedical systematic reviews, 16 percent of the included references found through databases were unique to one database(14). A particular four-source combination reached very high recall in that sample, but the result is not a universal prescription: it came from biomedical reviews at one institution, and specialized sources added unique records for some topics. The transferable finding is that source overlap is incomplete and must be examined rather than assumed.


Google Scholar illustrates another trap. Its breadth makes it valuable for supplementary discovery and citation tracing, but breadth is not the same as a reproducible systematic search. A prospective comparison covering 120 reviews found that Google Scholar, MEDLINE, and Embase were each insufficient when used alone(15). Google Scholar also limits practical access to a ranked subset of results, while its coverage and ranking rules are not exposed in the way a systematic search requires. The conclusion is not “never use Google Scholar.” It is “do not let one opaque ranking system define the evidence universe.”


A complete discovery system may combine subject databases, multidisciplinary citation indexes, regional sources, study registries, dissertation repositories, preprint servers that share manuscripts before formal journal review, government and organizational websites, conference records, and other grey literature. Grey literature means research-related material distributed outside conventional commercial journals, such as government reports, theses, technical assessments, and registry entries. It matters partly because publication is selective. If positive or statistically significant findings are more likely to become journal articles, a journal-only search can exaggerate apparent effects. Statistical significance indicates that an observed pattern would be unusual under a specified no-effect model; it does not by itself show that an effect is large or practically important. Tools such as Grey Matters turn website searching into a documented process by recording which sources were visited, which terms were used, and what was retrieved(20).


Supplementary methods cover failure modes that database queries cannot. Backward citation searching checks the references of a relevant paper. Forward citation searching identifies later papers that cite it. Related reviews can reveal older terminology and hard-to-find reports. Expert contact may locate ongoing or unpublished work. Handsearching can help when journals or conference proceedings are poorly indexed. The Campbell Collaboration's information-retrieval guide is especially useful outside medicine because it treats these routes as parts of a coherent plan rather than optional decoration(3).


A Better Search Strategy for Stronger Systematic Reviews


For complex questions, that plan may need several shorter, multi-stranded searches rather than one oversized query. NICE guidance derives the structure from the protocol and uses backward and forward chasing, handsearching, and expert contact when indexing is likely to miss evidence, while holding those supplementary methods to the same transparency and reproducibility expectations(6). NICE's pragmatic production limits are not defaults for every review; the transferable lesson is to specify every route before its results become tempting.


Searchers must also distinguish a study from a report. One underlying study may generate a registration record, protocol, conference abstract, primary article, follow-up analysis, and correction. Counting those reports as separate studies exaggerates the amount of independent evidence. Failing to link them can also split methods, outcomes, and harms across different documents. Search and record-management plans should therefore preserve report relationships from the beginning instead of attempting to repair them after synthesis.


Validate before launch and document during execution


Long search strategies often look authoritative, but length can hide defects. In an audit of 137 published PubMed or MEDLINE strategies, 92.7 percent contained at least one assessed error and 78.1 percent contained an error affecting recall(16). The study's 2018 sample and defined error taxonomy should not be generalized to every field or current interface. Even with that caveat, it demonstrates that published status and visual complexity are poor substitutes for quality control. Common failures included missing free-text or controlled terms, incomplete use of subject-heading hierarchies, and mistakes in truncation or phrase handling.


Independent peer review is one practical safeguard. The PRESS guideline uses six checks: whether the question was faithfully rendered; Boolean and proximity logic; controlled headings; free-text terms; spelling, syntax, and line construction; and imposed limits or filters(11). The review belongs after an initial database strategy exists but before that strategy is adapted for other platforms or run across the search plan. Corrections are then relatively cheap. PRESS does not judge the entire evidence-discovery plan, however. It cannot by itself determine whether the right databases were selected, whether grey literature matters, or whether supplementary searching and reporting are complete.


The detailed PRESS explanation and checklist places this review beside search planning, validation, and accurate reporting, and allows another review after a major revision(12). Peer review is therefore most useful as an early quality-control step, before a flawed strategy has created a large screening burden, rather than as a final endorsement of a finished search.


Validation also means examining what the strategy retrieves. Known-item testing asks whether the search finds seed studies. Term testing compares what controlled vocabulary retrieves with what free text retrieves, then inspects relevant records unique to either route. Result sampling checks whether irrelevant clusters reveal an ambiguous term that needs field restriction or whether a useful concept was made too narrow. The systematic approach described by Bramer and colleagues makes this iterative sequence explicit: clarify the question, identify concepts, collect subject headings and free-text terms, optimize, check errors, translate across databases, and test again(13).


Documentation must happen while the search is run. A final Boolean string alone is not enough. The record should include the database and platform, complete line-by-line strategy, date of execution, coverage dates where relevant, filters and limits, number of results, the method used to merge duplicate records retrieved from several sources, supplementary-search procedure, peer-review status, and update date. Platform matters because the same database can have different syntax and functions through different providers. Search date matters because database contents change.


That is easier to do contemporaneously than reconstruct after screening. In a historical audit of 65 Cochrane reviews, none reported all seven assessed search-description elements, with exact run dates and database hosts among the recurrent omissions(21). The 2006 sample predates PRISMA-S and does not estimate current practice, but it captures a persistent practical risk: details that were never recorded cannot be made reproducible by polishing the manuscript later.


PRISMA 2020 asks review authors to report the complete strategies used for all databases, registers, and websites rather than showing one illustrative query(7). Its detailed explanation and elaboration adds platforms, restrictions, filters, validation, peer review, and automated tools(8). PRISMA-S, a 16-item extension focused specifically on searching, covers sources, full strategies, citation searching, contacts, updates, dates, record counts, and deduplication(9). These are reporting standards, not certificates of methodological quality. A weak search can be reported perfectly. Their value is that a transparent record allows readers to reproduce the work, identify weaknesses, and update it later.


Updating should be planned rather than improvised. Months may pass between the original search and manuscript submission. New eligible studies can appear, indexing can change, and preprints can become journal articles. The protocol should state when the search will be rerun and how new records will enter screening. Even a rapid review does not escape the need for explicit tradeoffs. Guidance on rapid-review searching argues that abbreviations should be systematic, transparent, and reproducible(19). Reducing a justified source range is different from skipping quality assurance because time is short. Poor early design can actually increase total workload by flooding screening with avoidable noise or forcing a late rerun.


Search failures can be consequential, but a single criticism should not be converted into a prevalence estimate. In a contested critique and response, a librarian-revised MEDLINE strategy retrieved 1,016 additional records and identified omitted literature; the original authors accepted that their search could be improved while disputing whether the critics' three-part standard applied universally(18). The exchange is useful as a reminder to test systematic, comprehensive, and transparent practice, not as proof that every field will gain the same number of records.


The finished search is therefore not one impressive line of Boolean logic. It is a chain of accountable decisions: define the evidence boundary, translate only necessary concepts, discover and test vocabulary, choose complementary sources, trace citations and grey literature where appropriate, peer-review the main strategy, record every execution detail, and update before publication. When those decisions are made before the first production search, the team can explain both what it found and why the remaining uncertainty is acceptable. Without them, screening more records merely makes an undocumented search larger. It does not make the review systematic.


Sources

1. Cochrane Handbook, Chapter 4: Searching for and selecting studies - https://www.cochrane.org/authors/handbooks-and-manuals/handbook/current/chapter-04

2. JBI Manual for Evidence Synthesis - https://jbi-global-wiki.refined.site/space/MANUAL

3. Searching for studies: A guide to information retrieval for Campbell systematic reviews - https://doi.org/10.1002/cl2.1433

4. Standards for Finding and Assessing Individual Studies - https://www.ncbi.nlm.nih.gov/books/NBK209517/

5. Systematic Reviews: CRD's guidance for undertaking reviews in health care - https://www.york.ac.uk/media/crd/newpagesspring2025/Systematic_Reviews_CRD_guidance.pdf

6. NICE: Identifying the evidence - https://www.nice.org.uk/process/pmg20/chapter/identifying-the-evidence-literature-searching-and-evidence-submission

7. The PRISMA 2020 statement - https://doi.org/10.1371/journal.pmed.1003583

8. PRISMA 2020 explanation and elaboration - https://doi.org/10.1136/bmj.n160

9. PRISMA-S - https://doi.org/10.5195/jmla.2021.962

10. PRISMA-P 2015 statement - https://doi.org/10.1186/2046-4053-4-1

11. PRESS 2015 Guideline Statement - https://doi.org/10.1016/j.jclinepi.2016.01.021

12. PRESS 2015 Explanation and Elaboration - https://www.cda-amc.ca/sites/default/files/attachments/2023-06/PRESS Peer Review Electronic Search Strategies_ 2015 Guideline Explanation and Elaboration (PRESS E&E).pdf

13. A systematic approach to searching - https://doi.org/10.5195/jmla.2018.283

14. Optimal database combinations for systematic-review searches - https://doi.org/10.1186/s13643-017-0644-y

15. Comparing Embase, MEDLINE, and Google Scholar coverage - https://doi.org/10.1186/s13643-016-0215-7

16. Errors in systematic-review search strategies - https://doi.org/10.5195/jmla.2019.567

17. The impact of PICO as a search-strategy tool - https://doi.org/10.5195/jmla.2018.345

18. Systematic review searches must be systematic, comprehensive, and transparent - https://doi.org/10.1186/s12889-018-6275-y

19. Rapid reviews methods series: Guidance on literature search - https://doi.org/10.1136/bmjebm-2022-112079

20. Grey Matters - https://www.cda-amc.ca/node/88098

21. Analysis of search-strategy reporting in Cochrane reviews - https://doi.org/10.3163/1536-5050.97.1.004