Rainbow Padding: Mitigating Early Termination in Instruction-Tuned Diffusion LLMs
Bumjun Kim, Dongjae Jeon, Dueun Kim, Wonje Jeung, Albert No
Abstract
Diffusion large language models (dLLMs) have emerged as a promising alternative to autoregressive models, offering flexible generation orders and strong performance on complex reasoning tasks. However, instruction-tuned dLLMs exhibit a critical vulnerability we term <eos> overflow: as allocated sequence length increases, responses paradoxically become shorter, collapsing into early termination or degenerating into streams of <eos> tokens. Although noticed in practice, this issue has not been systematically analyzed. We trace its root cause to the dual role of <eos> as both termination and padding, which concentrates probability mass on <eos> at later positions and propagates backward to trigger early termination. To address this, we introduce Rainbow Padding, a simple remedy that replaces repeated <eos> placeholders with a repeating cycle of distinct padding tokens, distributing probability mass and breaking <eos> dominance. Experiments show that Rainbow Padding substantially improves length robustness and output quality, with as few as seven padding tokens sufficient to prevent early termination. Moreover, the method integrates efficiently into existing instruction-tuned models: LoRA fine-tuning for a single epoch on minimal data yields significant improvements, making this solution highly practical. The code is publicly available at https://github.com/quasar529/rainbow-padding . * Equal Contribution. † Corresponding Author. 1 Throughout this paper, LLaDA and Dream denote the instruction-tuned models LLaDA-8B-Instruct and Dream-v0-Instruct-7B, unless stated otherwise.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ebc3bd24-666c-417a-b07f-9bc6f991d323Cited by top-tier papers3
- DAPD: Dependency-Aware Parallel Decoding via Attention for Diffusion LLMsBumjun Kim, Dongjae Jeon, Moongyu Jeon, Albert NoICML 2026 · 8 citations
- Early Decisions Matter: Proximity Bias and Initial Trajectory Shaping in Non-Autoregressive Diffusion Language ModelsJiyeon Kim, Sungik Choi, Yongrae Jo, Moontae Lee et al.ICML 2026 · 4 citations
- Insertion Based Sequence Generation with Learnable Order DynamicsDhruvesh Patel, Benjamin Rozonoyer, Gaurav Pandey, Tahira Naseem et al.ICML 2026 · 1 citation
Builds on27
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
Related papers
- DPad: Efficient Diffusion Language Models with Suffix DropoutXinhua Chen, Sitao Huang, Cong Guo, Chiyue Wei et al.ICLR 2026 · 44 citations
- Residual Context Diffusion Language ModelsYuezhou Hu, Harman Singh, Monishwaran Maheswaran, Haocheng Xi et al.ICML 2026
- DreamOn: Diffusion Language Models For Code Infilling Beyond Fixed-size CanvasZirui Wu, Lin Zheng, Zhihui Xie, Jiacheng Ye et al.ICLR 2026 · 33 citations
- UltraLLaDA: Scaling the Context Length to 128K for Diffusion Large Language ModelsGuangxin He, Shen Nie, Fengqi Zhu, Yuankang Zhao et al.ICLR 2026 · 13 citations
- A2D: Any-Order, Any-Step Safety Alignment for Diffusion Language ModelsWonje Jeung, Sangyeon Yoon, Yoonjun Cho, Dongjae Jeon et al.ICLR 2026 · 10 citations
