Sticking to the Mean: Detecting Sticky Tokens in Text Embedding Models
Kexin Chen, Dongxia Wang, Yi Liu, Haonan Zhang, Wenhai Wang
摘要
Despite the widespread use of Transformerbased text embedding models in NLP tasks, surprising "sticky tokens" can undermine the reliability of embeddings. These tokens, when repeatedly inserted into sentences, pull sentence similarity toward a certain value, disrupting the normal distribution of embedding similarities and degrading downstream performance. In this paper, we systematically investigate such anomalous tokens, formally defining them and introducing an efficient detection method, Sticky Token Detector (STD), based on sentence and token filtering. Applying STD to 40 checkpoints across 14 model families, we discover a total of 868 sticky tokens. Our analysis reveals that these tokens often originate from special or unused entries in the vocabulary, as well as fragmented subwords from multilingual corpora. Notably, their presence does not strictly correlate with model size or vocabulary size. We further evaluate how sticky tokens affect downstream tasks like clustering and retrieval, observing substantial performance degradation that approaches 50% in certain cases. Through attention-layer analysis, we show that sticky tokens disproportionately dominate the model's internal representations, raising concerns about tokenization robustness. Our findings show the need for better tokenization strategies and model design to mitigate the impact of sticky tokens in future text embedding applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper7
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- On the Sentence Embeddings from Pre-trained Language ModelsBohan Li, Hao Zhou, Junxian He, Mingxuan Wang 等EMNLP 2020 · 被引用 538 次
- Improving Neural Language Generation with Spectrum ControlLingxiao Wang, Jing Huang, Kevin Huang, Ziniu Hu 等ICLR 2020 · 被引用 94 次
- Tokenization Is More Than CompressionCraig W. Schmidt, Varshini Reddy, Haoran Zhang, Alec Alameddine 等EMNLP 2024 · 被引用 16 次
相关 Paper
- Glitch Tokens in Large Language Models: Categorization Taxonomy and Effective DetectionYuxi Li, Yi Liu, Gelei Deng, Ying Zhang 等FSE 2024 · 被引用 12 次
- Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language ModelsSander Land, Max BartoloEMNLP 2024 · 被引用 4 次
- Why Mean Pooling Works: Quantifying Second-Order Collapse in Text EmbeddingsTomomasa Hara, Hiroto Kurita, Masaaki Imaizumi, Kentaro Inui 等ACL 2026 · 被引用 2 次
- LightToken: A Task and Model-agnostic Lightweight Token Embedding Framework for Pre-trained Language ModelsHaoyu Wang, Ruirui Li, Haoming Jiang, Zhengyang Wang 等KDD 2023 · 被引用 5 次
- Your UnEmbedding Matrix is Secretly a Feature Lens for Text EmbeddingsSonghao Wu, Zhongxin Chen, Yuxuan Liu, Heng Cui 等KDD 2026 · 被引用 1 次
