More Sail than Ballast: Addressing Harmful Knowledge Leakage in the Expansive Reasoning Space of LRMs
Qibing Ren, Xinhao Song, Ke Fan, Lijun Li, Zhanpeng Zhou, Gongshen Liu, Junchi Yan, Lizhuang Ma, Jing Shao
摘要
The capabilities of large language models (LLMs), particularly large reasoning models (LRMs), are rapidly advancing. This raises concerns about whether LRMs can maintain their safety awareness throughout long-form reasoning. Frustratingly, we identify a prevalent safety issue across LLMs and LRMs, where LRMs can reveal dangerous thoughts, leading to harmful knowledge elicitation when confronting sensitive yet benign topics. For example, when explaining the chemical context of Lewisite, a biological weapon, LRMs analyze its synthesis in their reasoning without recognizing the associated risks. We refer to this issue as the "unintended elicitation" issue. Experiments on our benchmark show that it is a common issue across current LRMs due to their strong multi-step reasoning capabilities. To address this issue, we propose placing LLMs in our synthesized open-ended environments, allowing them to self-search for a safety reasoning pattern to respond responsibly and helpfully. We first design a scalable data synthesis pipeline to generate data that triggers the unintended elicitation issue. We further propose a safety-first reward model design, which prioritizes safety while also evaluating the helpfulness of responses and the faithfulness of reasoning. Experiments show that our method improves safety, reduces overrefusal, and maintains strong helpfulness, paving the way for safer deployment in high-stakes domains. Code is available at https://github. com/XinhaoS0101/Safety-CoT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper27
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
相关 Paper
- Reasoning Model Unlearning: Forgetting Traces, Not Just Answers, While Preserving Reasoning SkillsChangsheng Wang, Chongyu Fan, Yihua Zhang, Jinghan Jia 等EMNLP 2025
- How Should We Enhance the Safety of Large Reasoning Models: An Empirical StudyZhexin Zhang, Xian Qi Loye, Victor Shea-Jay Huang, Junxiao Yang 等ACL 2026 · 被引用 20 次
- Finding and Reactivating Post-Trained LLMs' Hidden Safety MechanismsMingjie Li, Wai Man Si, Michael Backes, Yang Zhang 等NeurIPS 2025 · 被引用 4 次
- The CoT Encyclopedia: Analyzing, Predicting, and Controlling how a Reasoning Model will ThinkSeongyun Lee, Seungone Kim, Minju Seo, Yongrae Jo 等ICLR 2026 · 被引用 6 次
- SafeAgent: Safeguarding LLM Agents via an Automated Risk SimulatorXueyang Zhou, Weidong Wang, Lin Lu, Jiawen Shi 等ACL 2026 · 被引用 5 次
