A Text is Worth Several Tokens: Text Embedding from LLMs Secretly Aligns Well with The Key Tokens
Zhijie Nie, Richong Zhang, Zhanyu Wu
摘要
Text embeddings from large language models (LLMs) have achieved excellent results in tasks such as information retrieval, semantic textual similarity, etc. In this work, we show an interesting finding: when feeding a text into the LLM-based embedder, the obtained text embedding will be able to be aligned with the key tokens in the input text. We first fully analyze this phenomenon on eight LLM-based embedders and show that this phenomenon is universal and is not affected by model architecture, training strategy, and embedding method. With a deeper analysis, we find that the main change in embedding space between these embedders and their LLM backbones is in the first principal component. By adjusting the first principal component, we can align text embedding with the key tokens. Finally, we give several examples to demonstrate the vast application potential of this finding: (1) we propose a simple and practical sparse retrieval method based on the aligned tokens, which can achieve 80% of the dense retrieval effect of the same model while reducing the computation significantly; (2) we show that our findings provide a novel perspective to help understand novel technologies (e.g., instruction-following embedding) and fuzzy concepts (e.g., semantic relatedness vs. similarity) in this field.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Attention Illuminates LLM Reasoning: The Uncovered Preplan-and-Anchor Rhythm Enables Fine-Grained Policy OptimizationYang Li, Zhichen Dong, Yuhan Sun, Weixun Wang 等ICML 2026 · 被引用 25 次
- How Does Reasoning Flow? Tracing Attention-Induced Information Flow for Targeted RL in LLMsZhichen Dong, Yang Li, Yuhan Sun, Weixun Wang 等ICML 2026 · 被引用 1 次
- Your UnEmbedding Matrix is Secretly a Feature Lens for Text EmbeddingsSonghao Wu, Zhongxin Chen, Yuxuan Liu, Heng Cui 等KDD 2026 · 被引用 1 次
- InstEmb: Instruction-Following Embeddings through Glimpses of the FutureTianhao Gao, Jun Fang, Xiaohui Zhang, Zhiyuan Liu 等ICML 2026
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- Privacy Risks of General-Purpose Language ModelsXudong Pan, Mi Zhang, Shouling Ji, Min YangS&P 2020 · 被引用 291 次
- Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary SpaceMor Geva, Avi Caciularu, Kevin Ro Wang, Yoav GoldbergEMNLP 2022 · 被引用 92 次
相关 Paper
- Enhancing Lexicon-Based Text Embeddings with Large Language ModelsYibin Lei, Tao Shen, Yu Cao, Andrew YatesACL 2025
- Answer is All You Need: Instruction-following Text Embedding via Answering the QuestionLetian Peng, Yuwei Zhang, Zilong Wang, Jayanth Srinivasa 等ACL 2024
- Token Prepending: A Training-Free Approach for Eliciting Better Sentence Embeddings from LLMsYuchen Fu, Zifeng Cheng, Zhiwei Jiang, Zhonghui Wang 等ACL 2025
- Learning Retrieval Models with Sparse AutoencodersThibault Formal, Maxime Louis, Hervé Déjean, Stéphane ClinchantICLR 2026 · 被引用 9 次
- Large Language Models as Foundations for Next-Gen Dense Retrieval: A Comprehensive Empirical AssessmentKun Luo, Minghao Qin, Zheng Liu, Shitao Xiao 等EMNLP 2024 · 被引用 3 次
