Can ChatGPT Write a Good Boolean Query for Systematic Review Literature Search?
Shuai Wang, Harrisen Scells, Bevan Koopman, Guido Zuccon
摘要
BEVAN KOOPMAN, CSIRO, Australia Systematic reviews are comprehensive reviews of the literature for a highly focused research question. These reviews are often treated as the highest form of evidence in evidence-based medicine, and are the key strategy to answer research questions in the medical field. To create a high-quality systematic review, complex Boolean queries are often constructed to retrieve studies for the review topic. However, it often takes a long time for systematic review researchers to construct a high quality systematic review Boolean query, and often the resulting queries are far from effective. Poor queries may lead to biased or invalid reviews, because they missed to retrieve key evidence, or to extensive increase in review costs, because they retrieved too many irrelevant studies. Recent advances in Transformer-based generative models have shown great potential to effectively follow instructions from users and generate answers based on the instructions being made. In this paper, we investigate the effectiveness of the latest of such models, ChatGPT, in generating effective Boolean queries for systematic review literature search. Through a number of extensive experiments on standard test collections for the task, we find that ChatGPT is capable of generating queries that lead to high search precision, although trading-off this for recall. Overall, our study demonstrates the potential of ChatGPT in generating effective Boolean queries for systematic review literature search. The ability of ChatGPT to follow complex instructions and generate queries with high precision makes it a valuable tool for researchers conducting systematic reviews, particularly for rapid reviews where time is a constraint and often trading-off higher precision for lower recall is acceptable.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- In-Context Impersonation Reveals Large Language Models' Strengths and BiasesLeonard Salewski, Stephan Alaniz, Isabel Rio-Torto, Eric Schulz 等NeurIPS 2023 · 被引用 259 次
- A Setwise Approach for Effective and Highly Efficient Zero-shot Ranking with Large Language ModelsShengyao Zhuang, Honglei Zhuang, Bevan Koopman, Guido ZucconSIGIR 2024 · 被引用 60 次
- DiscipLink: Unfolding Interdisciplinary Information Seeking Process via Human-AI Co-ExplorationChengbo Zheng, Yuanhao Zhang, Zeyu Huang, Chuhan Shi 等UIST 2024 · 被引用 19 次
- Debate on Graph: A Flexible and Reliable Reasoning Framework for Large Language ModelsJie Ma, Zhitao Gao, Qi Chai, Wangchun Sun 等AAAI 2025 · 被引用 8 次
- Principled Personas: Defining and Measuring the Intended Effects of Persona Prompting on Task PerformancePedro Henrique Luz de Araujo, Paul Röttger, Dirk Hovy, Benjamin RothEMNLP 2025 · 被引用 1 次
它引用的顶会 Paper3
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Design Guidelines for Prompt Engineering Text-to-Image Generative ModelsVivian Liu, Lydia B. ChiltonCHI 2022 · 被引用 586 次
- Automatic Boolean Query Formulation for Systematic Review Literature SearchHarrisen Scells, Guido Zuccon, Bevan Koopman, Justin ClarkWWW 2020 · 被引用 51 次
相关 Paper
- Smooth Operators for Effective Systematic Review QueriesHarrisen Scells, Ferdinand Schlatt, Martin PotthastSIGIR 2023 · 被引用 4 次
- Appraising the Potential Uses and Harms of LLMs for Medical Systematic ReviewsHye Sun Yun, Iain James Marshall, Thomas A. Trikalinos, Byron C. WallaceEMNLP 2023 · 被引用 11 次
- Dr ChatGPT tell me what I want to hear: How different prompts impact health answer correctnessBevan Koopman, Guido ZucconEMNLP 2023 · 被引用 69 次
- Can Large Language Models Match the Conclusions of Systematic Reviews?Christopher Polzak, Alejandro Lozano, Min Woo Sun, James Burgess 等ICLR 2026 · 被引用 9 次
- Detecting Health Advice in Medical Research LiteratureYingya Li, Jun Wang, Bei YuEMNLP 2021 · 被引用 1 次
