Rare Tokens Degenerate All Tokens: Improving Neural Text Generation via Adaptive Gradient Gating for Rare Token Embeddings
Sangwon Yu, Jongyoon Song, Heeseung Kim, Seongmin Lee, Woo-Jong Ryu, Sungroh Yoon
Abstract
Recent studies have determined that the learned token embeddings of large-scale neural language models are degenerated to be anisotropic with a narrow-cone shape. This phenomenon, called the representation degeneration problem, facilitates an increase in the overall similarity between token embeddings that negatively affect the performance of the models. Although the existing methods that address the degeneration problem based on observations of the phenomenon triggered by the problem improves the performance of the text generation, the training dynamics of token embeddings behind the degeneration problem are still not explored. In this study, we analyze the training dynamics of the token embeddings focusing on rare token embedding. We demonstrate that the specific part of the gradient for rare token embeddings is the key cause of the degeneration problem for all tokens during training stage. Based on the analysis, we propose a novel method called, adaptive gradient gating (AGG). AGG addresses the degeneration problem by gating the specific part of the gradient for rare token embeddings. Experimental results from language modeling, word similarity, and machine translation tasks quantitatively and qualitatively verify the effectiveness of AGG.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 16900b5a-8036-49d1-acc9-94a1b60fdb4bCited by top-tier papers11
- CCTEST: Testing and Repairing Code Completion SystemsZongjie Li, Chaozheng Wang, Zhibo Liu, Haoxuan Wang et al.ICSE 2023 · 49 citations
- An Analysis of Tokenization: Transformers under Markov DataNived Rajaraman, Jiantao Jiao, Kannan RamchandranNeurIPS 2024 · 16 citations
- Metis: Training LLMs with FP4 QuantizationHengjie Cao, Mengyi Chen, Yifeng Yang, Fang Dong et al.ICLR 2026 · 10 citations
- Task-Adaptive Tokenization: Enhancing Long-Form Text Generation Efficacy in Mental Health and BeyondSiyang Liu, Naihao Deng, Sahand Sabour, Yilin Jia et al.EMNLP 2023 · 7 citations
- HyPe: Better Pre-trained Language Model Fine-tuning with Hidden Representation PerturbationHongyi Yuan, Zheng Yuan, Chuanqi Tan, Fei Huang et al.ACL 2023 · 5 citations
Builds on2
Related papers
- Straight to the Gradient: Learning to Use Novel Tokens for Neural Text GenerationXiang Lin, Simeng Han, Shafiq R. JotyICML 2021 · 30 citations
- A Contrastive Framework for Neural Text GenerationYixuan Su, Tian Lan, Yan Wang, Dani Yogatama et al.NeurIPS 2022 · 349 citations
- Repetition In Repetition Out: Towards Understanding Neural Text Degeneration from the Data PerspectiveHuayang Li, Tian Lan, Zihao Fu, Deng Cai et al.NeurIPS 2023 · 54 citations
- Sticking to the Mean: Detecting Sticky Tokens in Text Embedding ModelsKexin Chen, Dongxia Wang, Yi Liu, Haonan Zhang et al.ACL 2025
- One Token Can Help! Learning Scalable and Pluggable Virtual Tokens for Retrieval-Augmented Large Language ModelsYutao Zhu, Zhaoheng Huang, Zhicheng Dou, Ji-Rong WenAAAI 2025 · 9 citations
