RBCBF: Decoding Time Safety Alignment via Risk Guided Rollback and Barrier Control
Tianxiang Chen, Jingyuan Zhou, Longhao Yan, Kaidi Yang
摘要
Existing decoding-time safety interventions are often reactive, relying on local signals to correct unsafe outputs after they emerge. Under adversarial prompts that drive generation into recurring unsafe response, such local signals provide weak guidance for stable repair. As a result, rollback and post-hoc rewriting often trade-off response quality with recurrent violations. To address these limitations, we propose RBCBF, a rollback-based decoding-time framework that jointly selects intervention steps and performs distribution-level corrective control. Our key innovation is a riskaggregation formulation that views terminal violations as the accumulated build-up of risk along the prefix. By selecting rollback steps from these decisive prefixes, RBCBF moves rollback targeting beyond heuristic cues and turns it into a trajectorylevel decision. RBCBF then applies invasive corrective control to the next-token distribution under multiple rule constraints. Across jailbreakstyle evaluations, RBCBF outperforms prior rollback methods and decoding-time baselines, reducing harmful responses and substantially lowering violation recurrence. Code is available at https://github.com/XTERY11/RBCBF.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Large Language Models Cannot Self-Correct Reasoning YetJie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng 等ICLR 2024 · 被引用 858 次
- WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language ModelsLiwei Jiang, Kavel Rao, Seungju Han, Allyson Ettinger 等NeurIPS 2024 · 被引用 247 次
- Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank ModificationsBoyi Wei, Kaixuan Huang, Yangsibo Huang, Tinghao Xie 等ICML 2024 · 被引用 215 次
- Decoding-Time Language Model Alignment with Multiple ObjectivesRuizhe Shi, Yifang Chen, Yushi Hu, Alisa Liu 等NeurIPS 2024 · 被引用 111 次
- Decoding-time Realignment of Language ModelsTianlin Liu, Shangmin Guo, Leonardo Bianco, Daniele Calandriello 等ICML 2024 · 被引用 69 次
相关 Paper
- Reinforcement Learning with Backtracking FeedbackBilgehan Sel, Vaishakh Keshava, Phillip Wallis, Lukas Rutishauser 等NeurIPS 2025
- CARE: Decoding-Time Safety Alignment via Rollback and Introspection InterventionXiaomeng Hu, Fei Huang, Chenhan Yuan, Junyang Lin 等NeurIPS 2025 · 被引用 5 次
- SafeSpec: Fast and Safe LLM via Dynamic Reflective SamplingHAOTIAN XU, Zeyang Zhang, Linbao Li, Huadi Zheng 等ICML 2026
- Root Defense Strategies: Ensuring Safety of LLM at the Decoding LevelXinyi Zeng, Yuying Shang, Jiawei Chen, Jingyuan Zhang 等ACL 2025 · 被引用 6 次
- Path Drift in Large Reasoning Models: How First-Person Commitments Override SafetyYuyi Huang, Runzhe Zhan, Lidia S. Chao, Ailin Tao 等EMNLP 2025 · 被引用 3 次
