Your UnEmbedding Matrix is Secretly a Feature Lens for Text Embeddings
Songhao Wu, Zhongxin Chen, Yuxuan Liu, Heng Cui, Cong Li, Rui Yan
Abstract
Large language models exhibit impressive zero-shot capabilities across a wide range of downstream tasks. However, they struggle to function as off-the-shelf embedding models, leading to suboptimal performance on massive text embedding benchmarks. In this paper, we identify a potential cause underlying this deficiency. Our motivation stems from an unexpected observation: text embeddings tend to align with frequent but uninformative tokens when projected onto the vocabulary space. We argue that this excessive expression of high-frequency tokens suppresses the model's ability to capture nuanced semantics. To address this, we introduce EmbedFilter, a simple linear transformation designed to refine text embeddings derived from LLMs directly. Specifically, we uncover that the unembedding matrix within LLMs encodes a latent space that is actively writing these frequent tokens into embedding space. By filtering out this subspace, EmbedFilter suppress the influence of high-frequency tokens, thereby enhancing semantic representations. As a compelling byproduct, this enables an inherent dimensionality reduction, lowering index storage and speedup retrieval while fully preserving the refined embedding quality. Our experiments across multiple LLM backbones demonstrate that LLMs equipped with EmbedFilter achieve superior zero-shot downstream performance even with significantly reduced embedding dimensions. We hope our findings provide deeper insights into the mechanisms of LLM-based representations and inspire more principled designs to improve text embeddings training. Our code is available at https://github.com/CentreChen/EmbFilter. * These authors contributed equally. Songhao Wu discovered the core phenomenon, provided the core implementation and led the writing. Zhongxin Chen refined the code, conducted the experiments and provided Songhao Wu with valuable insights.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bde287fb-4f2f-4ec2-905d-601c9433a755Builds on10
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- On the Sentence Embeddings from Pre-trained Language ModelsBohan Li, Hao Zhou, Junxian He, Mingxuan Wang et al.EMNLP 2020 · 538 citations
- SPECTER: Document-level Representation Learning using Citation-informed TransformersArman Cohan, Sergey Feldman, Iz Beltagy, Doug Downey et al.ACL 2020 · 20 citations
- Fact or Fiction: Verifying Scientific ClaimsDavid Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang et al.EMNLP 2020 · 6 citations
- Meta-Task Prompting Elicits Embeddings from Large Language ModelsYibin Lei, Di Wu, Tianyi Zhou, Tao Shen et al.ACL 2024 · 6 citations
Related papers
- ScalingFilter: Assessing Data Quality through Inverse Utilization of Scaling LawsRuihang Li, Yixuan Wei, Miaosen Zhang, Nenghai Yu et al.EMNLP 2024 · 1 citation
- A Text is Worth Several Tokens: Text Embedding from LLMs Secretly Aligns Well with The Key TokensZhijie Nie, Richong Zhang, Zhanyu WuACL 2025 · 5 citations
- Enhancing Lexicon-Based Text Embeddings with Large Language ModelsYibin Lei, Tao Shen, Yu Cao, Andrew YatesACL 2025
- Matryoshka-Adaptor: Unsupervised and Supervised Tuning for Smaller Embedding DimensionsJinsung Yoon, Rajarishi Sinha, Sercan Ömer Arik, Tomas PfisterEMNLP 2024 · 1 citation
- LUSIFER: Language Universal Space Integration for Enhanced Representation in Multilingual Text Embedding ModelsHieu Man, Nghia Trung Ngo, Viet Dac Lai, Ryan A. Rossi et al.SIGIR 2025
