Resolving the Security-Auditability Dilemma with Auditable Latent Chain-of-Thought Alignment
Guan Wang, Biyu Zhou, Xuehai Tang, Jizhong Han, Songlin Hu
摘要
To address the increasingly severe safety risk of large language models (LLMs), reasoningbased safety alignment methods have emerged. These methods overcome the limitations of 'shallow alignment' by exposing the model's Chain-of-Thought (CoT), enabling auditability of safety reasoning process through both training-phase supervision and post-generation verification. However, this transparency creates a critical vulnerability, a tension we define as the Security Auditability Dilemma: while explicit reasoning is a prerequisite for safety, its textual Auditable paradoxically transforms it into an optimization target for adaptive attackers and induces the model to unintentionally copy harmful content from its own reasoning context. To address this, we propose Auditable Latent CoT Alignment (ALCA), a framework that decouples internal reasoning from external output. ALCA shifts the safety deliberation process into a continuous latent space. This allows the safety reasoning process to guide the generation of harmless outputs, while eliminates the discrete textual surface that facilitates internal copying and adaptive attack. Yet, this process is not a black box. we introduce a restricted Self-Decoding mechanism that allows the model to reconstruct its latent reasoning into human-readable text for supervision under specific guidance. Extensive experiments show that ALCA achieves robustness alignment, reducing the success rate of adaptive jailbreak attacks by over 40% compared to strong baselines, while preserving performance. Our framework presents a path toward building LLMs that are both robustly secure and auditable.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 被引用 722 次
相关 Paper
- More Thinking, Less Talking: Internalizing Deliberative Safety into LLM ParametersGuan Wang, Xuehai Tang, Biyu Zhou, Jizhong Han 等ACL 2026
- Alignment-Weighted DPO: A principled reasoning approach to improve safety alignmentMengxuan Hu, Vivek V. Datla, Anoop Kumar, Zihan Guan 等ICLR 2026 · 被引用 3 次
- AdvChain: Adversarial Chain-of-Thought Tuning for Robust Safety Alignment of Large Reasoning ModelsZihao Zhu, Xinyu Wu, Gehan Hu, Siwei Lyu 等ICLR 2026 · 被引用 6 次
- STAIR: Improving Safety Alignment with Introspective ReasoningYichi Zhang, Siyuan Zhang, Yao Huang, Zeyu Xia 等ICML 2025
- Safety Alignment Can Be Not Superficial With Explicit Safety SignalsJianwei Li, Jung-Eun KimICML 2025
