Making Large Language Models Efficient Dense Retrievers
Yibin Lei, Shwai He, Ang Li, Andrew Yates
摘要
Recent work has shown that directly fine-tuning large language models (LLMs) for dense retrieval yields strong performance, but their substantial parameter counts make them computationally inefficient. While prior studies have revealed significant layer redundancy in LLMs for generative tasks, it remains unclear whether similar redundancy exists when these models are adapted for retrieval tasks, which require encoding entire sequences into fixed representations rather than generating tokens iteratively. To this end, we conduct a comprehensive analysis of layer redundancy in LLM-based dense retrievers. We find that, in contrast to generative settings, MLP layers are substantially more prunable, while attention layers remain critical for semantic aggregation. Building on this insight, we propose EffiR, a framework for developing efficient retrievers that performs large-scale MLP compression through a coarseto-fine strategy (coarse-grained depth reduction followed by fine-grained width reduction), combined with retrieval-specific fine-tuning. Across diverse BEIR datasets and LLM backbones, EffiR achieves substantial reductions in model size and inference cost while preserving the performance of full-size models. 1 * Equal contribution 1 Our code and models are available at https://github. com/Yibin-Lei/EffiR .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper18
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang 等ICLR 2021 · 被引用 1,547 次
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 被引用 1,240 次
- Sheared LLaMA: Accelerating Language Model Pre-training via Structured PruningMengzhou Xia, Tianyu Gao, Zhiyuan Zeng, Danqi ChenICLR 2024 · 被引用 453 次
相关 Paper
- Large Language Models as Foundations for Next-Gen Dense Retrieval: A Comprehensive Empirical AssessmentKun Luo, Minghao Qin, Zheng Liu, Shitao Xiao 等EMNLP 2024 · 被引用 3 次
- DRAMA: Diverse Augmentation from Large Language Models to Smaller Dense RetrieversXueguang Ma, Xi Victoria Lin, Barlas Oguz, Jimmy Lin 等ACL 2025 · 被引用 20 次
- Efficient Sparse Retrieval with Lightweight Superblock PruningParker Carlson, Wentai Xie, Rohil Shah, Tao YangSIGIR 2026 · 被引用 2 次
- GRASP: Replace Redundant Layers with Adaptive Singular Parameters for Efficient Model CompressionKainan Liu, Yong Zhang, Ning Cheng, Zhitao Li 等EMNLP 2025
- EARN: Efficient Inference Acceleration for LLM-based Generative Recommendation by Register TokensChaoqun Yang, Xinyu Lin, Wenjie Wang, Yongqi Li 等KDD 2025 · 被引用 1 次
