Large Language Models as Foundations for Next-Gen Dense Retrieval: A Comprehensive Empirical Assessment
Kun Luo, Minghao Qin, Zheng Liu, Shitao Xiao, Jun Zhao, Kang Liu
摘要
Pre-trained language models like BERT and T5 serve as crucial backbone encoders for dense retrieval.However, these models often exhibit limited generalization capabilities and face challenges in improving in-domain accuracy.Recent research has explored using large language models (LLMs) as retrievers, achieving state-of-the-art performance across various tasks.Despite these advancements, the specific benefits of LLMs over traditional retrievers and the impact of different LLM configurations-such as parameter sizes, pre-training duration, and alignment processes-on retrieval tasks remain unclear.In this work, we conduct a comprehensive empirical study on six key dimensions of dense retrieval capabilities, including in-domain accuracy, data efficiency, zero-shot generalization, lengthy retrieval, instruction-based retrieval, and multi-task learning.We evaluate over 15 different backbone LLMs and non-LLMs.Our findings reveal that larger models and extensive pre-training consistently enhance in-domain accuracy and data efficiency.Additionally, larger models demonstrate significant potential in zero-shot generalization, lengthy retrieval, instruction-based retrieval, and multi-task learning.These results underscore the advantages of LLMs as versatile and effective backbone encoders in dense retrieval, providing valuable insights for future research and development in this field.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- DRAMA: Diverse Augmentation from Large Language Models to Smaller Dense RetrieversXueguang Ma, Xi Victoria Lin, Barlas Oguz, Jimmy Lin 等ACL 2025 · 被引用 20 次
- CoQuIR: A Comprehensive Benchmark for Code Quality-Aware Information RetrievalJiahui Geng, Fengyu Cai, Shaobo Cui, Qing Li 等ACL 2026 · 被引用 3 次
- Making Large Language Models Efficient Dense RetrieversYibin Lei, Shwai He, Ang Li, Andrew YatesACL 2026 · 被引用 2 次
- Causal2Vec: Improving Decoder-only LLMs as Embedding Models through a Contextual TokenAiliang Lin, Zhuoyun Li, Yusong Wang, Kotaro Funakoshi 等ACL 2026
- SURE or Not? Investigating Semantic Understanding in Dense Retrieval ModelsLingdi Kong, Xuanang Chen, Ben He, Le SunACL 2026
它引用的顶会 Paper11
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang 等ICLR 2021 · 被引用 1,547 次
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIsYujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu 等ICLR 2024 · 被引用 1,469 次
- Large Dual Encoders Are Generalizable RetrieversJianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai 等EMNLP 2022 · 被引用 145 次
相关 Paper
- Scaling Retrieval-Based Language Models with a Trillion-Token DatastoreRulin Shao, Jacqueline He, Akari Asai, Weijia Shi 等NeurIPS 2024 · 被引用 76 次
- ChatRetriever: Adapting Large Language Models for Generalized and Robust Conversational Dense RetrievalKelong Mao, Chenlong Deng, Haonan Chen, Fengran Mo 等EMNLP 2024 · 被引用 7 次
- Evaluating the Effectiveness and Scalability of LLM-Based Data Augmentation for RetrievalPranjal A. Chitale, Bishal Santra, Yashoteja Prabhu, Amit SharmaEMNLP 2025
- Llama2Vec: Unsupervised Adaptation of Large Language Models for Dense RetrievalChaofan Li, Zheng Liu, Shitao Xiao, Yingxia Shao 等ACL 2024 · 被引用 10 次
- ReLLa: Retrieval-enhanced Large Language Models for Lifelong Sequential Behavior Comprehension in RecommendationJianghao Lin, Rong Shan, Chenxu Zhu, Kounianhua Du 等WWW 2024 · 被引用 151 次
