A Theoretical Analysis of the Repetition Problem in Text Generation
Zihao Fu, Wai Lam, Anthony Man-Cho So, Bei Shi
摘要
Text generation tasks, including translation, summarization, language models, and etc. see rapid growth during recent years. Despite the remarkable achievements, the repetition problem has been observed in nearly all text generation models undermining the generation performance extensively. To solve the repetition problem, many methods have been proposed, but there is no existing theoretical analysis to show why this problem happens and how it is resolved. In this paper, we propose a new framework for theoretical analysis for the repetition problem. We first define the Average Repetition Probability (ARP) to characterize the repetition problem quantitatively. Then, we conduct an extensive analysis of the Markov generation model and derive several upper bounds of the average repetition probability with intuitive understanding. We show that most of the existing methods are essentially minimizing the upper bounds explicitly or implicitly. Grounded on our theory, we show that the repetition problem is, unfortunately, caused by the traits of our language itself. One major reason is attributed to the fact that there exist too many words predicting the same word as the subsequent word with high probability. Consequently, it is easy to go back to that word and form repetitions and we dub it as the high inflow problem. Furthermore, we extend our analysis to broader generation models by deriving a concentration bound of the average repetition probability for a general generation model. Finally, based on the theoretical upper bounds, we propose a novel rebalanced encoding approach to alleviate the high inflow problem and thus reducing the upper bound. The experimental results show that our theoretical framework is applicable in general generation models and our proposed rebalanced encoding approach alleviates the repetition problem significantly in both the translation task and the language modeling task. The source code of this paper can be obtained from https://github.com/fuzihaofzh/repetition-problem-nlg.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper31
- On the Effectiveness of Parameter-Efficient Fine-TuningZihao Fu, Haoran Yang, Anthony Man-Cho So, Wai Lam 等AAAI 2023 · 被引用 234 次
- Learning to Break the Loop: Analyzing and Mitigating Repetitions for Neural Text GenerationJin Xu, Xiaojiang Liu, Jianhao Yan, Deng Cai 等NeurIPS 2022 · 被引用 135 次
- Stay on Topic with Classifier-Free GuidanceGuillaume Sanchez, Alexander Spangher, Honglu Fan, Elad Levi 等ICML 2024 · 被引用 76 次
- Repetition In Repetition Out: Towards Understanding Neural Text Degeneration from the Data PerspectiveHuayang Li, Tian Lan, Zihao Fu, Deng Cai 等NeurIPS 2023 · 被引用 54 次
- From Self-Attention to Markov Models: Unveiling the Dynamics of Generative TransformersMuhammed Emrullah Ildiz, Yixiao Huang, Yingcong Li, Ankit Singh Rawat 等ICML 2024 · 被引用 45 次
它引用的顶会 Paper6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan 等ICLR 2020 · 被引用 683 次
- Language GANs Falling ShortMassimo Caccia, Lucas Caccia, William Fedus, Hugo Larochelle 等ICLR 2020 · 被引用 236 次
- Partially-Aligned Data-to-Text Generation with Distant SupervisionZihao Fu, Bei Shi, Wai Lam, Lidong Bing 等EMNLP 2020 · 被引用 19 次
相关 Paper
- Rethinking Repetition Problems of LLMs in Code GenerationYihong Dong, Yuchen Liu, Xue Jiang, Bin Gu 等ACL 2025
- Context Tokens are Anchors: Understanding the Repeat Curse in dMLLMs from an Information Flow PerspectiveQiyan Zhao, Xiaofeng Zhang, Shuochen Chang, Qianyu Chen 等ICLR 2026 · 被引用 3 次
- CoNT: Contrastive Neural Text GenerationChenxin An, Jiangtao Feng, Kai Lv, Lingpeng Kong 等NeurIPS 2022 · 被引用 37 次
- Language Generation with Replay: A Learning-Theoretic View of Model CollapseGiorgio Racca, Michal Valko, Amartya SanyalICML 2026 · 被引用 4 次
- Rare Tokens Degenerate All Tokens: Improving Neural Text Generation via Adaptive Gradient Gating for Rare Token EmbeddingsSangwon Yu, Jongyoon Song, Heeseung Kim, Seongmin Lee 等ACL 2022 · 被引用 42 次
