On the Theoretical Limitations of Embedding-Based Retrieval
Orion Weller, Michael Boratko, Iftekhar Naim, Jinhyuk Lee
摘要
Vector embeddings have been tasked with an ever-increasing set of retrieval tasks over the years, with a nascent rise in using them for reasoning, instruction-following, coding, and more. These new benchmarks push embeddings to work for any query and any notion of relevance that could be given. While prior works have pointed out theoretical limitations of vector embeddings, there is a common assumption that these difficulties are exclusively due to unrealistic queries, and those that are not can be overcome with better training data and larger models. In this work, we demonstrate that we may encounter these theoretical limitations in realistic settings with extremely simple queries. We connect known results in learning theory, showing that the number of top-k subsets of documents capable of being returned as the result of some query is limited by the dimension of the embedding. We empirically show that this holds true even if we directly optimize on the test set with free parameterized embeddings. Using free embeddings, we then demonstrate that returning all pairs of documents requires a relatively high dimension. We then create a realistic dataset called LIMIT that stress tests embedding models based on these theoretical results, and observe that even state-of-the-art models fail on this dataset despite the simple nature of the task. Our work shows the limits of embedding models under the existing single vector paradigm and calls for future research to develop new techniques that can resolve this fundamental limitation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late InteractionZilin Xiao, Qi Ma, Mengting Gu, Chun-cheng Jason Chen 等ICLR 2026 · 被引用 40 次
- The Geometry of Reasoning: Flowing Logics in Representation SpaceYufa Zhou, Yixiao Wang, Xunjian Yin, Shuyan Zhou 等ICLR 2026 · 被引用 29 次
- Understanding Generative Recommendation with Semantic IDs from a Model-scaling ViewJingzhe Liu, Liam Collins, Jiliang Tang, Tong Zhao 等KDD 2026 · 被引用 17 次
- MILCO: Learned Sparse Retrieval Across Languages via a Multilingual ConnectorThong Nguyen, Yibin Lei, Jia-Huei Ju, Eugene Yang 等ICLR 2026 · 被引用 16 次
- Featurized-Decomposition Join: Low-Cost Semantic Joins with GuaranteesSepanta Zeighami, Shreya Shankar, Aditya G. ParameswaranVLDB 2026 · 被引用 11 次
它引用的顶会 Paper15
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 被引用 1,246 次
- Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and InferenceBenjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller 等ACL 2025 · 被引用 552 次
- Matryoshka Representation LearningAditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford 等NeurIPS 2022 · 被引用 364 次
- Seq vs Seq: An Open Suite of Paired Encoders and DecodersOrion Weller, Kathryn Ricci, Marc Marone, Antoine Chaffin 等ICLR 2026 · 被引用 50 次
相关 Paper
- Scaling Laws for Embedding Dimension in Information RetrievalJulian Killingback, Mahta Rafiee, Madine Manas, Hamed ZamaniSIGIR 2026
- RMIR: A Benchmark Dataset for Reasoning-Intensive Multimodal Image RetrievalYijiang Li, Kunal Kotian, Ali Marjaninejad, Meir Friedenberg 等CVPR 2026
- BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive RetrievalHongjin Su, Howard Yen, Mengzhou Xia, Weijia Shi 等ICLR 2025
- is Theoretically Large Enough for Embedding-based Top- RetrievalZihao Wang, Hang Yin, Lihui Liu, Hanghang Tong 等ICML 2026
- Provable Accuracy Collapse of Embedding-Based Representations under Dimensionality MismatchDionysis Arvanitakis, Vaggos Chatziafratis, Yiyuan LuoICML 2026
