To Generate or to Retrieve? On the Effectiveness of Artificial Contexts for Medical Open-Domain Question Answering
Giacomo Frisoni, Alessio Cocchieri, Alex Presepi, Gianluca Moro, Zaiqiao Meng
Abstract
Medical open-domain question answering demands substantial access to specialized knowledge. Recent efforts have sought to decouple knowledge from model parameters, counteracting architectural scaling and allowing for training on common low-resource hardware. The retrieve-then-read paradigm has become ubiquitous, with model predictions grounded on relevant knowledge pieces from external repositories such as PubMed, textbooks, and UMLS. An alternative path, still under-explored but made possible by the advent of domain-specific large language models, entails constructing artificial contexts through prompting. As a result, "to generate or to retrieve" is the modern equivalent of Hamlet's dilemma. This paper presents MEDGENIE, the first generate-thenread framework for multiple-choice question answering in medicine. We conduct extensive experiments on MedQA-USMLE, MedMCQA, and MMLU, incorporating a practical perspective by assuming a maximum of 24GB VRAM. MEDGENIE sets a new state-of-the-art in the open-book setting of each testbed, allowing a small-scale reader to outcompete zero-shot closed-book 175B baselines while using up to 706× fewer parameters. Our findings reveal that generated passages are more effective than retrieved ones in attaining higher accuracy. 1 * Equal contribution (co-first authorship). 1 Our code, fine-tuned models, and generated contexts are publicly available at https://github.com/unibo-nlp/medgenie .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d73c9ff6-e5ed-4554-9fa9-13805dacc5dbCited by top-tier papers10
- Unknown Claims: Generation of Fact-Checking Training Examples from Unstructured and Structured DataJean-Flavien Bussotti, Luca Ragazzi, Giacomo Frisoni, Gianluca Moro et al.EMNLP 2024 · 3 citations
- Sycophants in the Courtroom: Are LLMs Fragile to Juridical Authority and Evolving Legal Standards?Lorenzo Molfetta, Alessio Cocchieri, Luca Ragazzi, Ilaria Bartolini et al.ACL 2026 · 1 citation
- Can Large Language Models Win the International Mathematical Games?Alessio Cocchieri, Luca Ragazzi, Giuseppe Tagliavini, Lorenzo Tordi et al.EMNLP 2025
- LLM-Based Multi-Agent Systems for Clinical Workflows: A Survey of AI HospitalsZonghai Yao, Hong YuACL 2026
- Time-MQA: Time Series Multi-Task Question Answering with Context EnhancementYaxuan Kong, Yiyuan Yang, Yoontae Hwang, Wenjie Du et al.ACL 2025
Builds on11
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong et al.ICML 2022 · 1,173 citations
- Memorization Without Overfitting: Analyzing the Training Dynamics of Large Language ModelsKushal Tirumala, Aram H. Markosyan, Luke Zettlemoyer, Armen AghajanyanNeurIPS 2022 · 304 citations
- Deep Bidirectional Language-Knowledge Graph PretrainingMichihiro Yasunaga, Antoine Bosselut, Hongyu Ren, Xikun Zhang et al.NeurIPS 2022 · 294 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
Related papers
- From Retrieval to Generation: Unifying External and Parametric Knowledge for Medical Question AnsweringLei Li, Xiao Zhou, Yingying Zhang, Xian WuWWW 2026
- Generate rather than Retrieve: Large Language Models are Strong Context GeneratorsWenhao Yu, Dan Iter, Shuohang Wang, Yichong Xu et al.ICLR 2023 · 86 citations
- Harnessing Multi-Role Capabilities of Large Language Models for Open-Domain Question AnsweringHongda Sun, Yuxuan Liu, Chengwei Wu, Haiyu Yan et al.WWW 2024 · 16 citations
- Generate-then-Ground in Retrieval-Augmented Generation for Multi-hop Question AnsweringZhengliang Shi, Shuo Zhang, Weiwei Sun, Shen Gao et al.ACL 2024
- Merging Generated and Retrieved Knowledge for Open-Domain QAYunxiang Zhang, Muhammad Khalifa, Lajanugen Logeswaran, Moontae Lee et al.EMNLP 2023 · 12 citations
