Rainbow Padding: Mitigating Early Termination in Instruction-Tuned Diffusion LLMs
Bumjun Kim, Dongjae Jeon, Dueun Kim, Wonje Jeung, Albert No
摘要
Diffusion large language models (dLLMs) have emerged as a promising alternative to autoregressive models, offering flexible generation orders and strong performance on complex reasoning tasks. However, instruction-tuned dLLMs exhibit a critical vulnerability we term <eos> overflow: as allocated sequence length increases, responses paradoxically become shorter, collapsing into early termination or degenerating into streams of <eos> tokens. Although noticed in practice, this issue has not been systematically analyzed. We trace its root cause to the dual role of <eos> as both termination and padding, which concentrates probability mass on <eos> at later positions and propagates backward to trigger early termination. To address this, we introduce Rainbow Padding, a simple remedy that replaces repeated <eos> placeholders with a repeating cycle of distinct padding tokens, distributing probability mass and breaking <eos> dominance. Experiments show that Rainbow Padding substantially improves length robustness and output quality, with as few as seven padding tokens sufficient to prevent early termination. Moreover, the method integrates efficiently into existing instruction-tuned models: LoRA fine-tuning for a single epoch on minimal data yields significant improvements, making this solution highly practical. The code is publicly available at https://github.com/quasar529/rainbow-padding . * Equal Contribution. † Corresponding Author. 1 Throughout this paper, LLaDA and Dream denote the instruction-tuned models LLaDA-8B-Instruct and Dream-v0-Instruct-7B, unless stated otherwise.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- DAPD: Dependency-Aware Parallel Decoding via Attention for Diffusion LLMsBumjun Kim, Dongjae Jeon, Moongyu Jeon, Albert NoICML 2026 · 被引用 8 次
- Early Decisions Matter: Proximity Bias and Initial Trajectory Shaping in Non-Autoregressive Diffusion Language ModelsJiyeon Kim, Sungik Choi, Yongrae Jo, Moontae Lee 等ICML 2026 · 被引用 4 次
- Insertion Based Sequence Generation with Learnable Order DynamicsDhruvesh Patel, Benjamin Rozonoyer, Gaurav Pandey, Tahira Naseem 等ICML 2026 · 被引用 1 次
它引用的顶会 Paper27
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
相关 Paper
- DPad: Efficient Diffusion Language Models with Suffix DropoutXinhua Chen, Sitao Huang, Cong Guo, Chiyue Wei 等ICLR 2026 · 被引用 44 次
- Residual Context Diffusion Language ModelsYuezhou Hu, Harman Singh, Monishwaran Maheswaran, Haocheng Xi 等ICML 2026
- DreamOn: Diffusion Language Models For Code Infilling Beyond Fixed-size CanvasZirui Wu, Lin Zheng, Zhihui Xie, Jiacheng Ye 等ICLR 2026 · 被引用 33 次
- UltraLLaDA: Scaling the Context Length to 128K for Diffusion Large Language ModelsGuangxin He, Shen Nie, Fengqi Zhu, Yuankang Zhao 等ICLR 2026 · 被引用 13 次
- A2D: Any-Order, Any-Step Safety Alignment for Diffusion Language ModelsWonje Jeung, Sangyeon Yoon, Yoonjun Cho, Dongjae Jeon 等ICLR 2026 · 被引用 10 次
