RBCBF: Decoding Time Safety Alignment via Risk Guided Rollback and Barrier Control
Tianxiang Chen, Jingyuan Zhou, Longhao Yan, Kaidi Yang
Abstract
Existing decoding-time safety interventions are often reactive, relying on local signals to correct unsafe outputs after they emerge. Under adversarial prompts that drive generation into recurring unsafe response, such local signals provide weak guidance for stable repair. As a result, rollback and post-hoc rewriting often trade-off response quality with recurrent violations. To address these limitations, we propose RBCBF, a rollback-based decoding-time framework that jointly selects intervention steps and performs distribution-level corrective control. Our key innovation is a riskaggregation formulation that views terminal violations as the accumulated build-up of risk along the prefix. By selecting rollback steps from these decisive prefixes, RBCBF moves rollback targeting beyond heuristic cues and turns it into a trajectorylevel decision. RBCBF then applies invasive corrective control to the next-token distribution under multiple rule constraints. Across jailbreakstyle evaluations, RBCBF outperforms prior rollback methods and decoding-time baselines, reducing harmful responses and substantially lowering violation recurrence. Code is available at https://github.com/XTERY11/RBCBF.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cfc07cd1-4856-4fa3-9366-c716cedc94a6Builds on13
- Large Language Models Cannot Self-Correct Reasoning YetJie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng et al.ICLR 2024 · 858 citations
- WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language ModelsLiwei Jiang, Kavel Rao, Seungju Han, Allyson Ettinger et al.NeurIPS 2024 · 247 citations
- Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank ModificationsBoyi Wei, Kaixuan Huang, Yangsibo Huang, Tinghao Xie et al.ICML 2024 · 215 citations
- Decoding-Time Language Model Alignment with Multiple ObjectivesRuizhe Shi, Yifang Chen, Yushi Hu, Alisa Liu et al.NeurIPS 2024 · 111 citations
- Decoding-time Realignment of Language ModelsTianlin Liu, Shangmin Guo, Leonardo Bianco, Daniele Calandriello et al.ICML 2024 · 69 citations
Related papers
- Reinforcement Learning with Backtracking FeedbackBilgehan Sel, Vaishakh Keshava, Phillip Wallis, Lukas Rutishauser et al.NeurIPS 2025
- CARE: Decoding-Time Safety Alignment via Rollback and Introspection InterventionXiaomeng Hu, Fei Huang, Chenhan Yuan, Junyang Lin et al.NeurIPS 2025 · 5 citations
- SafeSpec: Fast and Safe LLM via Dynamic Reflective SamplingHAOTIAN XU, Zeyang Zhang, Linbao Li, Huadi Zheng et al.ICML 2026
- Root Defense Strategies: Ensuring Safety of LLM at the Decoding LevelXinyi Zeng, Yuying Shang, Jiawei Chen, Jingyuan Zhang et al.ACL 2025 · 6 citations
- Path Drift in Large Reasoning Models: How First-Person Commitments Override SafetyYuyi Huang, Runzhe Zhan, Lidia S. Chao, Ailin Tao et al.EMNLP 2025 · 3 citations
