A2D: Any-Order, Any-Step Safety Alignment for Diffusion Language Models
Wonje Jeung, Sangyeon Yoon, Yoonjun Cho, Dongjae Jeon, Sangwoo Shin, Hyesoo Hong, Albert No
摘要
Diffusion large language models (dLLMs) enable any-order generation, but this flexibility enlarges the attack surface: harmful spans may appear at arbitrary positions, and template-based prefilling attacks such as DIJA bypass response-level refusals. We introduce A2D (Any-Order, Any-Step Defense), a token-level alignment method that aligns dLLMs to emit an [EOS] refusal signal whenever harmful content arises. By aligning safety directly at the token-level under randomized masking, A2D achieves robustness to both any-decoding-order and any-step prefilling attacks under various conditions. It also enables real-time monitoring: dLLMs may begin a response but automatically terminate if unsafe continuation emerges. On safety benchmarks, A2D consistently prevents the generation of harmful outputs, slashing DIJA success rates from over 80% to near-zero (1.3% on LLaDA-8B-Instruct, 0.0% on Dream-v0-Instruct-7B), and thresholded [EOS] probabilities allow early rejection, yielding up to 19.3× faster safe termination.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper32
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
相关 Paper
- The Devil behind the mask: An emergent safety vulnerability of Diffusion LLMsZichen Wen, Jiashu Qu, Zhaorun Chen, Xiaoya Lu 等ICLR 2026 · 被引用 32 次
- Any-Depth Alignment: Unlocking Innate Safety Alignment of LLMs to Any-DepthJiawei Zhang, Andrew Estornell, David D. Baek, Bo Li 等ICLR 2026 · 被引用 3 次
- Where to Start Alignment? Diffusion Large Language Model May Demand a Distinct PositionZhixin Xie, Xurui Song, Jun LuoAAAI 2026 · 被引用 7 次
- Improving LLM Safety Alignment with Dual-Objective OptimizationXuandong Zhao, Will Cai, Tianneng Shi, David Huang 等ICML 2025
- Toward Safer Diffusion Language Models: Discovery and Mitigation of Priming VulnerabilityShojiro Yamabe, Jun SakumaICLR 2026 · 被引用 9 次
