On the Theoretical Limitations of Embedding-Based Retrieval
Orion Weller, Michael Boratko, Iftekhar Naim, Jinhyuk Lee
Abstract
Vector embeddings have been tasked with an ever-increasing set of retrieval tasks over the years, with a nascent rise in using them for reasoning, instruction-following, coding, and more. These new benchmarks push embeddings to work for any query and any notion of relevance that could be given. While prior works have pointed out theoretical limitations of vector embeddings, there is a common assumption that these difficulties are exclusively due to unrealistic queries, and those that are not can be overcome with better training data and larger models. In this work, we demonstrate that we may encounter these theoretical limitations in realistic settings with extremely simple queries. We connect known results in learning theory, showing that the number of top-k subsets of documents capable of being returned as the result of some query is limited by the dimension of the embedding. We empirically show that this holds true even if we directly optimize on the test set with free parameterized embeddings. Using free embeddings, we then demonstrate that returning all pairs of documents requires a relatively high dimension. We then create a realistic dataset called LIMIT that stress tests embedding models based on these theoretical results, and observe that even state-of-the-art models fail on this dataset despite the simple nature of the task. Our work shows the limits of embedding models under the existing single vector paradigm and calls for future research to develop new techniques that can resolve this fundamental limitation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5832a89a-a105-419c-8c09-dffad1ba73a2Cited by top-tier papers23
- MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late InteractionZilin Xiao, Qi Ma, Mengting Gu, Chun-cheng Jason Chen et al.ICLR 2026 · 40 citations
- The Geometry of Reasoning: Flowing Logics in Representation SpaceYufa Zhou, Yixiao Wang, Xunjian Yin, Shuyan Zhou et al.ICLR 2026 · 29 citations
- Understanding Generative Recommendation with Semantic IDs from a Model-scaling ViewJingzhe Liu, Liam Collins, Jiliang Tang, Tong Zhao et al.KDD 2026 · 17 citations
- MILCO: Learned Sparse Retrieval Across Languages via a Multilingual ConnectorThong Nguyen, Yibin Lei, Jia-Huei Ju, Eugene Yang et al.ICLR 2026 · 16 citations
- Featurized-Decomposition Join: Low-Cost Semantic Joins with GuaranteesSepanta Zeighami, Shreya Shankar, Aditya G. ParameswaranVLDB 2026 · 11 citations
Builds on15
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and InferenceBenjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller et al.ACL 2025 · 552 citations
- Matryoshka Representation LearningAditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford et al.NeurIPS 2022 · 364 citations
- Seq vs Seq: An Open Suite of Paired Encoders and DecodersOrion Weller, Kathryn Ricci, Marc Marone, Antoine Chaffin et al.ICLR 2026 · 50 citations
Related papers
- Scaling Laws for Embedding Dimension in Information RetrievalJulian Killingback, Mahta Rafiee, Madine Manas, Hamed ZamaniSIGIR 2026
- RMIR: A Benchmark Dataset for Reasoning-Intensive Multimodal Image RetrievalYijiang Li, Kunal Kotian, Ali Marjaninejad, Meir Friedenberg et al.CVPR 2026
- BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive RetrievalHongjin Su, Howard Yen, Mengzhou Xia, Weijia Shi et al.ICLR 2025
- is Theoretically Large Enough for Embedding-based Top- RetrievalZihao Wang, Hang Yin, Lihui Liu, Hanghang Tong et al.ICML 2026
- Provable Accuracy Collapse of Embedding-Based Representations under Dimensionality MismatchDionysis Arvanitakis, Vaggos Chatziafratis, Yiyuan LuoICML 2026
