LitSearch: A Retrieval Benchmark for Scientific Literature Search
Anirudh Ajith, Mengzhou Xia, Alexis Chevalier, Tanya Goyal, Danqi Chen, Tianyu Gao
Abstract
Literature search questions, such as "Where can I find research on the evaluation of consistency in generated summaries?" pose significant challenges for modern search engines and retrieval systems. These questions often require a deep understanding of research concepts and the ability to reason across entire articles. In this work, we introduce LitSearch, a retrieval benchmark comprising 597 realistic literature search queries about recent ML and NLP papers. Lit-Search is constructed using a combination of (1) questions generated by GPT-4 based on paragraphs containing inline citations from research papers and (2) questions manually written by authors about their recently published papers. All LitSearch questions were manually examined or edited by experts to ensure high quality. We extensively benchmark state-ofthe-art retrieval models and also evaluate two LLM-based reranking pipelines. We find a significant performance gap between BM25 and state-of-the-art dense retrievers, with a 24.8% absolute difference in recall@5. The LLMbased reranking strategies further improve the best-performing dense retriever by 4.4%. Additionally, commercial search engines and research tools like Google Search perform poorly on LitSearch, lagging behind the best dense retriever by up to 32 recall points. Taken together, these results show that LitSearch is an informative new testbed for retrieval systems while catering to a real-world use case. 1 Author-written Question: Invite ACL'23/ICLR'24 authors to write a question for their own papers Inline-citation Question: Sample an inline citation and prompt GPT-4 to write a question (Figure 2) Target Paper Target Paper Which method involves training additional prompt tokens for every layer during the fine-tuning of language models? Can you find a research paper that uses structured pruning techniques to scale down language models, where the original model being pruned has billions of parameters?
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers21
- AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research SuiteJonathan Bragg, Mike D'Arcy, Nishant Balepur, Dan Bareket et al.ICLR 2026 · 51 citations
- Can LLMs Identify Critical Limitations within Scientific Research? A Systematic Evaluation on AI Research PapersZhijian Xu, Yilun Zhao, Manasi Patwardhan, Lovekesh Vig et al.ACL 2025 · 17 citations
- Literature Meets Data: A Synergistic Approach to Hypothesis GenerationHaokun Liu, Yangqiaoyu Zhou, Mingxuan Li, Chenfei Yuan et al.ACL 2025 · 17 citations
- AgenticScholar: Agentic Data Management with Pipeline Orchestration for Scholarly CorporaHai Lan, Tingting Wang, Zhifeng Bao, Guoliang Li et al.SIGMOD 2026 · 4 citations
- XtraGPT: Context-Aware and Controllable Academic Paper Revision via Human-AI CollaborationNuo Chen, Andre Huikai Lin, Jiaying Wu, Junyi Hou et al.ACL 2026 · 3 citations
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- S2ORC: The Semantic Scholar Open Research CorpusKyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney et al.ACL 2020 · 424 citations
- Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking AgentsWeiwei Sun, Lingyong Yan, Xinyu Ma, Shuaiqiang Wang et al.EMNLP 2023 · 182 citations
Related papers
- Reinforced Query Reasoners for Reasoning-intensive Retrieval TasksXubo Qin, Jun Bai, Jiaqi Li, Zixia Jia et al.EMNLP 2025
- LiveNewsBench: Evaluating Web Search Agents with Freshly Curated NewsYunfan Zhang, Kathleen McKeown, Smaranda MuresanICML 2026 · 2 citations
- LitFM: A Retrieval Augmented Structure-aware Foundation Model For Citation GraphsJiasheng Zhang, Ali Maatouk, Jialin Chen, Ngoc Bui et al.KDD 2025 · 2 citations
- From Missteps to Mastery: Enhancing Low-Resource Dense Retrieval through Adaptive Query GenerationZhenyu Tong, Chuan Qin, Chuyu Fang, Kaichun Yao et al.KDD 2025 · 4 citations
- IntrAgent: An LLM Agent for Content-Grounded Information Retrieval through Literature ReviewFengbo Ma, Zixin Rao, Xiaoting Li, Zhetao Chen et al.ACL 2026 · 2 citations
