CARE: Decoding-Time Safety Alignment via Rollback and Introspection Intervention
Xiaomeng Hu, Fei Huang, Chenhan Yuan, Junyang Lin, Tsung-Yi Ho
摘要
As large language models (LLMs) are increasingly deployed in real-world applications, ensuring the safety of their outputs during decoding has become a critical challenge. However, existing decoding-time interventions, such as Contrastive Decoding, often force a severe trade-off between safety and response quality. In this work, we propose CARE, a novel framework for decoding-time safety alignment that integrates three key components: (1) a guard model for real-time safety monitoring, enabling detection of potentially unsafe content; (2) a rollback mechanism with a token buffer to correct unsafe outputs efficiently at an earlier stage without disrupting the user experience; and (3) a novel introspection-based intervention strategy, where the model generates self-reflective critiques of its previous outputs and incorporates these reflections into the context to guide subsequent decoding steps. The framework achieves a superior safety-quality trade-off by using its guard model for precise interventions, its rollback mechanism for timely corrections, and our novel introspection method for effective self-correction. Experimental results demonstrate that our framework achieves a superior balance of safety, quality, and efficiency, attaining a low harmful response rate and minimal disruption to the user experience while maintaining high response quality.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper9
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Tree of Attacks: Jailbreaking Black-Box LLMs AutomaticallyAnay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson 等NeurIPS 2024 · 被引用 835 次
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 被引用 722 次
- Many-shot JailbreakingCem Anil, Esin Durmus, Nina Panickssery, Mrinank Sharma 等NeurIPS 2024 · 被引用 338 次
- ARGS: Alignment as Reward-Guided SearchMaxim Khanov, Jirayu Burapacheep, Yixuan LiICLR 2024 · 被引用 101 次
相关 Paper
- Safety Alignment of Large Language Models via Contrasting Safe and Harmful DistributionsXiaoyun Zhang, Zhengyue Zhao, Wenxuan Shi, Kaidi Xu 等AAAI 2026 · 被引用 4 次
- Root Defense Strategies: Ensuring Safety of LLM at the Decoding LevelXinyi Zeng, Yuying Shang, Jiawei Chen, Jingyuan Zhang 等ACL 2025 · 被引用 6 次
- Navigating the OverKill in Large Language ModelsChenyu Shi, Xiao Wang, Qiming Ge, Songyang Gao 等ACL 2024
- C-SafeGen: Certified Safe LLM Generation with Claim-Based Streaming GuardrailsMintong Kang, Zhaorun Chen, Bo LiNeurIPS 2025 · 被引用 4 次
- Please refuse to answer me! Mitigating Over-Refusal in Large Language Models via Adaptive Contrastive DecodingYupeng Qi, Ziyu Lyu, Lixin Cui, Lu Bai 等ACL 2026
