Randomly Removing 50% of Dimensions in Text Embeddings has Minimal Impact on Retrieval and Classification Tasks
Sotaro Takeshita, Yurina Takeshita, Daniel Ruffinelli, Simone Paolo Ponzetto
摘要
In this paper, we study the surprising impact that truncating text embeddings has on downstream performance. We consistently observe across 6 state-of-the-art text encoders and 26 downstream tasks, that randomly removing up to 50% of embedding dimensions results in only a minor drop in performance, less than 10%, in retrieval and classification tasks. Given the benefits of using smaller-sized embeddings, as well as the potential insights about text encoding, we study this phenomenon and find that, contrary to what is suggested in prior work, this is not the result of an ineffective use of representation space. Instead, we find that a large number of uniformly distributed dimensions actually cause an increase in performance when removed. This would explain why, on average, removing a large number of embedding dimensions results in a marginal drop in performance. We make similar observations when truncating the embeddings used by large language models to make next-token predictions on generative tasks, suggesting that this phenomenon is not isolated to classification or retrieval tasks. Our code is attached to the submission. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Why Mean Pooling Works: Quantifying Second-Order Collapse in Text EmbeddingsTomomasa Hara, Hiroto Kurita, Masaaki Imaizumi, Kentaro Inui 等ACL 2026 · 被引用 2 次
- Provable Accuracy Collapse of Embedding-Based Representations under Dimensionality MismatchDionysis Arvanitakis, Vaggos Chatziafratis, Yiyuan LuoICML 2026
它引用的顶会 Paper21
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 被引用 2,360 次
- Understanding Dimensional Collapse in Contrastive Self-supervised LearningLi Jing, Pascal Vincent, Yann LeCun, Yuandong TianICLR 2022 · 被引用 467 次
- Matryoshka Representation LearningAditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford 等NeurIPS 2022 · 被引用 364 次
相关 Paper
- Your UnEmbedding Matrix is Secretly a Feature Lens for Text EmbeddingsSonghao Wu, Zhongxin Chen, Yuxuan Liu, Heng Cui 等KDD 2026 · 被引用 1 次
- Layer by Layer: Uncovering Hidden Representations in Language ModelsOscar Skean, Md Rifat Arefin, Dan Zhao, Niket Patel 等ICML 2025
- Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference ModelsZachary Ankner, Cody Blakeney, Kartik Sreenivasan, Max Marion 等ICLR 2025 · 被引用 4 次
- A Text is Worth Several Tokens: Text Embedding from LLMs Secretly Aligns Well with The Key TokensZhijie Nie, Richong Zhang, Zhanyu WuACL 2025 · 被引用 5 次
- From Fully Trained to Fully Random Embeddings: Improving Neural Machine Translation with Compact Word Embedding TablesKrtin Kumar, Peyman Passban, Mehdi Rezagholizadeh, Yiu Sing Lau 等AAAI 2022 · 被引用 3 次
