LongLLaDA: Unlocking Long Context Capabilities in Diffusion LLMs
Xiaoran Liu, Yuerong Song, Zhigeng Liu, Zengfeng Huang, Qipeng Guo, Ziwei He, Xipeng Qiu
摘要
Large Language Diffusion Models, or diffusion LLMs, have emerged as a significant focus in NLP research, with substantial effort directed toward understanding their scalability and downstream task performance. However, their long-context capabilities remain unexplored, lacking systematic analysis or methods for context extension. In this work, we present the first systematic investigation comparing the long-context performance of diffusion LLMs and traditional auto-regressive LLMs. We first identify a unique characteristic of diffusion LLMs, unlike autoregressive LLMs, they maintain remarkably stable perplexity during direct context extrapolation. Moreover, where auto-regressive models fail outright during the Needle-In-A-Haystack task with context exceeding their pretrained length, we discover diffusion LLMs exhibit a distinct "local perception" phenomenon, enabling successful retrieval from recent context segments. We explain both phenomena through the lens of Rotary Position Embedding (RoPE) scaling theory. Building on these observations, we propose LongLLaDA, a training-free method that integrates LLaDA with the NTK-based RoPE extrapolation. Our results validate that established extrapolation scaling laws remain effective for extending the context windows of diffusion LLMs. Furthermore, we identify long-context tasks where diffusion LLMs outperform auto-regressive LLMs and others where they fall short. Consequently, this study establishes the first length extrapolation method for diffusion LLMs while providing essential theoretical insights and empirical benchmarks critical for advancing future research on long-context diffusion LLMs. The code is available at https://github.com/OpenMOSS/LongLLaDA .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Masks Can Be Distracting: On Context Comprehension in Diffusion Language ModelsJulianna Piskorz, Cristina Pinneri, Alvaro Correia, Motasem Alfarra 等ICML 2026 · 被引用 5 次
- Adversarial Reinforcement Learning for Robust Diffusion Large Language Model UnlearningZhiwei Zhang, Yudi Lin, Linlin Wu, Fali Wang 等ICML 2026
它引用的顶会 Paper23
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 被引用 1,168 次
- Large Language Diffusion ModelsShen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang 等NeurIPS 2025 · 被引用 949 次
- Simple and Effective Masked Diffusion Language ModelsSubham S. Sahoo, Marianne Arriola, Yair Schiff, Aaron Gokaslan 等NeurIPS 2024 · 被引用 929 次
- Faith and Fate: Limits of Transformers on CompositionalityNouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li 等NeurIPS 2023 · 被引用 728 次
相关 Paper
- UltraLLaDA: Scaling the Context Length to 128K for Diffusion Large Language ModelsGuangxin He, Shen Nie, Fengqi Zhu, Yuankang Zhao 等ICLR 2026 · 被引用 13 次
- Scaling Laws of RoPE-based ExtrapolationXiaoran Liu, Hang Yan, Chenxin An, Xipeng Qiu 等ICLR 2024 · 被引用 130 次
- Base of RoPE Bounds Context LengthMingyu Xu, Xin Men, Bingning Wang, Qingyu Zhang 等NeurIPS 2024 · 被引用 56 次
- Extending LLM Context Window with Adaptive Grouped Positional Encoding: A Training-Free MethodXinhao Xu, Jiaxin Li, Hui Chen, Zijia Lin 等ACL 2025 · 被引用 1 次
- CLEX: Continuous Length Extrapolation for Large Language ModelsGuanzheng Chen, Xin Li, Zaiqiao Meng, Shangsong Liang 等ICLR 2024 · 被引用 39 次
