BioReader: a Retrieval-Enhanced Text-to-Text Transformer for Biomedical Literature
Giacomo Frisoni, Miki Mizutani, Gianluca Moro, Lorenzo Valgimigli
Abstract
The latest batch of research has equipped language models with the ability to attend over relevant and factual information from nonparametric external sources, drawing a complementary path to architectural scaling. Besides mastering language, exploiting and contextualizing the latent world knowledge is crucial in complex domains like biomedicine. However, most works in the field rely on general-purpose models supported by databases like Wikipedia and Books. We introduce BIOREADER 1 , the first retrieval-enhanced text-to-text model for biomedical natural language processing. Our domain-specific T5-based solution augments the input prompt by fetching and assembling relevant scientific literature chunks from a neural database with ≈60 million tokens centered on PubMed. We fine-tune and evaluate BIORE-ADER on a broad array of downstream tasks, significantly outperforming several state-of-theart methods despite using up to 3x fewer parameters. In tandem with extensive ablation studies, we show that domain knowledge can be easily altered or supplemented to make the model generate correct predictions bypassing the retraining step and thus addressing the literature overload issue.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c72f2f82-b367-4d03-9eef-e02b0b89697bCited by top-tier papers6
- A Textbook Remedy for Domain Shifts: Knowledge Priors for Medical Image AnalysisYue Yang, Mona Gandhi, Yufei Wang, Yifan Wu et al.NeurIPS 2024 · 21 citations
- BMRetriever: Tuning Large Language Models as Better Biomedical Text RetrieversRan Xu, Wenqi Shi, Yue Yu, Yuchen Zhuang et al.EMNLP 2024 · 8 citations
- To Generate or to Retrieve? On the Effectiveness of Artificial Contexts for Medical Open-Domain Question AnsweringGiacomo Frisoni, Alessio Cocchieri, Alex Presepi, Gianluca Moro et al.ACL 2024 · 8 citations
- ReFusion: Improving Natural Language Understanding with Computation-Efficient Retrieval Representation FusionShangyu Wu, Ying Xiong, Yufei Cui, Xue Liu et al.ICLR 2024 · 7 citations
- Unknown Claims: Generation of Fact-Checking Training Examples from Unstructured and Structured DataJean-Flavien Bussotti, Luca Ragazzi, Giacomo Frisoni, Gianluca Moro et al.EMNLP 2024 · 3 citations
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai et al.ICML 2022 · 1,629 citations
Related papers
- BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific LiteratureAlejandro Lozano, Min Woo Sun, James Burgess, Liangyu Chen et al.CVPR 2025
- Retrieval-Augmented Multimodal Language ModelingMichihiro Yasunaga, Armen Aghajanyan, Weijia Shi, Richard James et al.ICML 2023 · 153 citations
- Improving Biomedical Abstractive Summarisation with Knowledge Aggregation from Citation PapersChen Tang, Shun Wang, Tomas Goldsack, Chenghua LinEMNLP 2023 · 5 citations
- LitFM: A Retrieval Augmented Structure-aware Foundation Model For Citation GraphsJiasheng Zhang, Ali Maatouk, Jialin Chen, Ngoc Bui et al.KDD 2025 · 2 citations
- BioTool: A Comprehensive Tool-Calling Dataset for Enhancing Biomedical Capabilities of Large Language ModelsXin Gao, Ruiyi Zhang, Meixi Du, Peijia Qin et al.ACL 2026
