Where to Start Alignment? Diffusion Large Language Model May Demand a Distinct Position
Zhixin Xie, Xurui Song, Jun Luo
Abstract
Diffusion Large Language Models (dLLMs) have recently emerged as a competitive non-autoregressive paradigm due to their unique training and inference approach. However, there is currently a lack of safety study on this novel architecture. In this paper, we present the first analysis of dLLMs' safety performance and propose a novel safety alignment method tailored to their unique generation characteristics. Specifically, we identify a critical asymmetry between the defender and attacker in terms of security. For the defender, we reveal that the middle tokens of the response, rather than the initial ones, are more critical to the overall safety of dLLM outputs; this seems to suggest that aligning middle tokens can be more beneficial to the defender. The attacker, on the contrary, may have limited power to manipulate middle tokens, as we find dLLMs have a strong tendency towards a sequential generation order in practice, forcing the attack to meet this distribution and diverting it from influencing the critical middle tokens. Building on this asymmetry, we introduce Middle-tOken Safety Alignment (MOSA), a novel method that directly aligns the model's middle generation with safe refusals exploiting reinforcement learning. We implement MOSA and compare its security performance against eight attack methods on two benchmarks. We also test the utility of MOSA-aligned dLLM on coding, math, and general reasoning. The results strongly prove the superiority of MOSA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4c773080-9200-4b0d-b003-3be10b25b906Cited by top-tier papers4
- DiffuGuard: How Intrinsic Safety is Lost and Found in Diffusion Large Language ModelsZherui Li, Zheng Nie, Zhenhong Zhou, Yue Liu et al.ICLR 2026 · 12 citations
- A2D: Any-Order, Any-Step Safety Alignment for Diffusion Language ModelsWonje Jeung, Sangyeon Yoon, Yoonjun Cho, Dongjae Jeon et al.ICLR 2026 · 10 citations
- Toward Safer Diffusion Language Models: Discovery and Mitigation of Priming VulnerabilityShojiro Yamabe, Jun SakumaICLR 2026 · 9 citations
- Every Step Counts: Decoding Trajectories as Authorship Fingerprints of dLLMsQi Li, Runpeng Yu, Haiquan Lu, Xinchao WangICML 2026 · 4 citations
Builds on22
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow et al.NeurIPS 2021 · 2,256 citations
- Jailbroken: How Does LLM Safety Training Fail?Alexander Wei, Nika Haghtalab, Jacob SteinhardtNeurIPS 2023 · 2,230 citations
Related papers
- The Devil behind the mask: An emergent safety vulnerability of Diffusion LLMsZichen Wen, Jiashu Qu, Zhaorun Chen, Xiaoya Lu et al.ICLR 2026 · 32 citations
- Any-Depth Alignment: Unlocking Innate Safety Alignment of LLMs to Any-DepthJiawei Zhang, Andrew Estornell, David D. Baek, Bo Li et al.ICLR 2026 · 3 citations
- Safety Alignment Can Be Not Superficial With Explicit Safety SignalsJianwei Li, Jung-Eun KimICML 2025
- Reinforcement Learning for Diffusion LLMs via Energy-Based Gibbs AlignmentYijia Fan, Jing Yang, Mingyu Liu, Kaitong Cai et al.ACL 2026
- Safety Alignment of LMs via Non-cooperative GamesAnselm Paulus, Ilia Kulikov, Brandon Amos, REMI MUNOS et al.ICML 2026 · 4 citations
