Adaptive Sparsity Optimization with Learnable Soft Top-K and Per-Term Thresholding for Efficient Retrieval
Wentai Xie, Parker Carlson, Shanxiu He, Tao Yang
Abstract
Recent work on neural sparse retrieval has demonstrated strong relevance by leveraging Large Language Models (LLMs) for semantic term expansion. However, learned models paired with previous sparsification techniques still yield overly long document and query vectors partly due to a large LLM vocabulary, imposing a serious challenge to retrieval time and space efficiency. This paper proposes a scheme for optimizing model sparsity through a synergy of adaptive strategies, including learnable soft top-??, per-term thresholding, and FLOPs regularization to increase the sparsity of query and document vectors. Experimental results with Lion-SP model on the MS MARCO and BEIR datasets demonstrate that the proposed scheme can outperform the baselines by significantly reducing the average query and document lengths. Our scheme can achieve much shorter retrieval latency and lower storage cost while maintaining highly competitive relevance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 576c8ce2-ccaa-4ad2-8e46-5d2896bbdd59Builds on9
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- Minimizing FLOPs to Learn Efficient Sparse RepresentationsBiswajit Paria, Chih-Kuan Yeh, Ian En-Hsu Yen, Ning Xu et al.ICLR 2020 · 85 citations
- Efficient Inverted Indexes for Approximate Retrieval over Learned Sparse RepresentationsSebastian Bruch, Franco Maria Nardini, Cosimo Rulli, Rossano VenturiniSIGIR 2024 · 47 citations
- LexMAE: Lexicon-Bottlenecked Pretraining for Large-Scale RetrievalTao Shen, Xiubo Geng, Chongyang Tao, Can Xu et al.ICLR 2023 · 14 citations
Related papers
- Efficient Sparse Retrieval with Lightweight Superblock PruningParker Carlson, Wentai Xie, Rohil Shah, Tao YangSIGIR 2026 · 2 citations
- Ultra-High Dimensional Sparse Representations with Binarization for Efficient Text RetrievalKyoungrok Jang, Junmo Kang, Giwon Hong, Sung-Hyon Myaeng et al.EMNLP 2021 · 13 citations
- Accelerating Inference of Retrieval-Augmented Generation via Sparse Context SelectionYun Zhu, Jia-Chen Gu, Caitlin Sikora, Ho Ko et al.ICLR 2025
- Learning Retrieval Models with Sparse AutoencodersThibault Formal, Maxime Louis, Hervé Déjean, Stéphane ClinchantICLR 2026 · 9 citations
- Making Large Language Models Efficient Dense RetrieversYibin Lei, Shwai He, Ang Li, Andrew YatesACL 2026 · 2 citations
