Can ChatGPT Write a Good Boolean Query for Systematic Review Literature Search?
Shuai Wang, Harrisen Scells, Bevan Koopman, Guido Zuccon
Abstract
BEVAN KOOPMAN, CSIRO, Australia Systematic reviews are comprehensive reviews of the literature for a highly focused research question. These reviews are often treated as the highest form of evidence in evidence-based medicine, and are the key strategy to answer research questions in the medical field. To create a high-quality systematic review, complex Boolean queries are often constructed to retrieve studies for the review topic. However, it often takes a long time for systematic review researchers to construct a high quality systematic review Boolean query, and often the resulting queries are far from effective. Poor queries may lead to biased or invalid reviews, because they missed to retrieve key evidence, or to extensive increase in review costs, because they retrieved too many irrelevant studies. Recent advances in Transformer-based generative models have shown great potential to effectively follow instructions from users and generate answers based on the instructions being made. In this paper, we investigate the effectiveness of the latest of such models, ChatGPT, in generating effective Boolean queries for systematic review literature search. Through a number of extensive experiments on standard test collections for the task, we find that ChatGPT is capable of generating queries that lead to high search precision, although trading-off this for recall. Overall, our study demonstrates the potential of ChatGPT in generating effective Boolean queries for systematic review literature search. The ability of ChatGPT to follow complex instructions and generate queries with high precision makes it a valuable tool for researchers conducting systematic reviews, particularly for rapid reviews where time is a constraint and often trading-off higher precision for lower recall is acceptable.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 810f857b-4db5-43d6-867b-62122239cdbaCited by top-tier papers8
- In-Context Impersonation Reveals Large Language Models' Strengths and BiasesLeonard Salewski, Stephan Alaniz, Isabel Rio-Torto, Eric Schulz et al.NeurIPS 2023 · 259 citations
- A Setwise Approach for Effective and Highly Efficient Zero-shot Ranking with Large Language ModelsShengyao Zhuang, Honglei Zhuang, Bevan Koopman, Guido ZucconSIGIR 2024 · 60 citations
- DiscipLink: Unfolding Interdisciplinary Information Seeking Process via Human-AI Co-ExplorationChengbo Zheng, Yuanhao Zhang, Zeyu Huang, Chuhan Shi et al.UIST 2024 · 19 citations
- Debate on Graph: A Flexible and Reliable Reasoning Framework for Large Language ModelsJie Ma, Zhitao Gao, Qi Chai, Wangchun Sun et al.AAAI 2025 · 8 citations
- Principled Personas: Defining and Measuring the Intended Effects of Persona Prompting on Task PerformancePedro Henrique Luz de Araujo, Paul Röttger, Dirk Hovy, Benjamin RothEMNLP 2025 · 1 citation
Builds on3
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Design Guidelines for Prompt Engineering Text-to-Image Generative ModelsVivian Liu, Lydia B. ChiltonCHI 2022 · 586 citations
- Automatic Boolean Query Formulation for Systematic Review Literature SearchHarrisen Scells, Guido Zuccon, Bevan Koopman, Justin ClarkWWW 2020 · 51 citations
Related papers
- Smooth Operators for Effective Systematic Review QueriesHarrisen Scells, Ferdinand Schlatt, Martin PotthastSIGIR 2023 · 4 citations
- Appraising the Potential Uses and Harms of LLMs for Medical Systematic ReviewsHye Sun Yun, Iain James Marshall, Thomas A. Trikalinos, Byron C. WallaceEMNLP 2023 · 11 citations
- Dr ChatGPT tell me what I want to hear: How different prompts impact health answer correctnessBevan Koopman, Guido ZucconEMNLP 2023 · 69 citations
- Can Large Language Models Match the Conclusions of Systematic Reviews?Christopher Polzak, Alejandro Lozano, Min Woo Sun, James Burgess et al.ICLR 2026 · 9 citations
- Detecting Health Advice in Medical Research LiteratureYingya Li, Jun Wang, Bei YuEMNLP 2021 · 1 citation
