Retrieval with Learned Similarities
Bailu Ding, Jiaqi Zhai
Abstract
Retrieval plays a fundamental role in recommendation systems, search, and natural language processing (NLP) by efficiently finding relevant items from a large corpus given a query. Dot products have been widely used as the similarity function in such tasks, enabled by Maximum Inner Product Search (MIPS) algorithms for efficient retrieval. However, state-of-the-art retrieval algorithms have migrated to learned similarities. These advanced approaches encompass multiple query embeddings, complex neural networks, direct item ID decoding via beam search, and hybrid solutions. Unfortunately, we lack efficient solutions for retrieval in these state-of-theart setups. Our work addresses this gap by investigating efficient retrieval techniques with expressive learned similarity functions. We establish Mixture-of-Logits (MoL) as a universal approximator of similarity functions, demonstrate that MoL's expressiveness can be realized empirically to achieve superior performance on diverse retrieval scenarios, and propose techniques to retrieve the approximate top-𝑘 results using MoL with tight error bounds. Through extensive experimentation, we show that MoL, enhanced by our proposed mutual information-based load balancing loss, sets new state-of-the-art results across heterogeneous scenarios, including sequential retrieval models in recommendation systems and finetuning language models for question answering; and our approximate top-𝑘 algorithms outperform baselines by up to 66× in latency while achieving > .99 recall rate compared to exact algorithms. 1 CCS Concepts
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3acf7f88-ec50-4fb5-82fc-1a5a16bb7a01Cited by top-tier papers3
- MINT: Multi-Vector Search Index TuningJiongli Zhu, Yue Wang, Bailu Ding, Philip A. Bernstein et al.ICDE 2026 · 1 citation
- SilverTorch: A Unified Model-based System to Democratize Large-Scale Recommendation on GPUsBi Xue, Hong Wu, Lei Chen, Chao Yang et al.SIGIR 2026
- Learning to Curate Context: Jointly Optimizing Retrieval and Prediction for Multimodal Social Media PopularityXovee Xu, Shuojun Lin, Fan Zhou, Jingkuan SongAAAI 2026
Builds on17
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai et al.ICML 2022 · 1,629 citations
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- Accelerating Large-Scale Inference with Anisotropic Vector QuantizationRuiqi Guo, Philip Sun, Erik Lindgren, Quan Geng et al.ICML 2020 · 539 citations
- Transformer Memory as a Differentiable Search IndexYi Tay, Vinh Tran, Mostafa Dehghani, Jianmo Ni et al.NeurIPS 2022 · 506 citations
Related papers
- Query-Aware Quantization for Maximum Inner Product SearchJin Zhang, Defu Lian, Haodi Zhang, Baoyun Wang et al.AAAI 2023 · 15 citations
- Anisotropic Additive Quantization for Fast Inner Product SearchJin Zhang, Qi Liu, Defu Lian, Zheng Liu et al.AAAI 2022 · 12 citations
- Relevance-Based Embeddings: Lightweight Candidate Retrieval via Heavy-Ranker CallsKirill Shevkunov, Andrey Ploskonosov, Liudmila ProkhorenkovaICML 2026
- Norm Adjusted Proximity Graph for Fast Inner Product RetrievalShulong Tan, Zhaozhuo Xu, Weijie Zhao, Hongliang Fei et al.KDD 2021 · 20 citations
- Multivariate Representation Learning for Information RetrievalHamed Zamani, Michael BenderskySIGIR 2023 · 7 citations
