Scaling Laws For Dense Retrieval
Yan Fang, Jingtao Zhan, Qingyao Ai, Jiaxin Mao, Weihang Su, Jia Chen, Yiqun Liu
Abstract
Scaling laws have been observed in a wide range of tasks, particularly in language generation. Previous studies have found that the performance of large language models adheres to predictable patterns with respect to the size of models and datasets. This helps us design training strategies effectively and efficiently, especially as large-scale training becomes increasingly resource-intensive. Yet, in dense retrieval, such scaling law has not been fully explored. In this study, we investigate how scaling affects the performance of dense retrieval models. We implement dense retrieval models with different numbers of parameters, and train them with various amounts of annotated data. We propose to use the contrastive entropy as the evaluation metric, which is continuous compared with discrete ranking metrics and thus can accurately reflect model performance. Results indicate that the performance of dense retrieval models follows a precise power-law scaling related to the model size and the number of annotations across different datasets and annotation methods. Additionally, we show that the scaling laws help optimize the training process, such as resolving the resource allocation problem under a budget constraint. We believe that these findings significantly contribute to understanding the scaling effect of dense retrieval models and offer meaningful guidance for future research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 36f5722c-fe11-4ff7-9582-a141cebf0751Cited by top-tier papers19
- PromptReps: Prompting Large Language Models to Generate Dense and Sparse Representations for Zero-Shot Document RetrievalShengyao Zhuang, Xueguang Ma, Bevan Koopman, Jimmy Lin et al.EMNLP 2024 · 26 citations
- Parametric Retrieval Augmented GenerationWeihang Su, Yichen Tang, Qingyao Ai, Junxi Yan et al.SIGIR 2025 · 25 citations
- Joint Evaluation of Answer and Reasoning Consistency for Hallucination Detection in Large Reasoning ModelsChangyue Wang, Weihang Su, Qingyao Ai, Yiqun LiuAAAI 2026 · 13 citations
- Best-of-Majority: Minimax-Optimal Strategy for Pass@k Inference ScalingQiwei Di, Kaixuan Ji, Xuheng Li, Heyang Zhao et al.ICLR 2026 · 7 citations
- ExpandR: Teaching Dense Retrievers Beyond Queries with LLM GuidanceSijia Yao, Pengcheng Huang, Zhenghao Liu, Yu Gu et al.EMNLP 2025 · 6 citations
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
Related papers
- Exploring Training and Inference Scaling Laws in Generative RetrievalHongru Cai, Yongqi Li, Ruifeng Yuan, Wenjie Wang et al.SIGIR 2025 · 1 citation
- Scaling Laws for Embedding Dimension in Information RetrievalJulian Killingback, Mahta Rafiee, Madine Manas, Hamed ZamaniSIGIR 2026
- On the Scaling of Robustness and Effectiveness in Dense RetrievalYu-An Liu, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke et al.SIGIR 2025 · 1 citation
- Inference Scaling Law for Retrieval Augmented GenerationShu Zhou, Yuxuan Ao, Yunyang Xuan, Xin Wang et al.AAAI 2026 · 1 citation
- Large Language Models as Foundations for Next-Gen Dense Retrieval: A Comprehensive Empirical AssessmentKun Luo, Minghao Qin, Zheng Liu, Shitao Xiao et al.EMNLP 2024 · 3 citations
