CoT is Not the Chain of Truth: An Empirical Internal Analysis of Reasoning LLMs for Fake News Generation
Zhao Tong, Chunlin Gong, Yiping Zhang, Haichao Shi, Qiang Liu, Xingcheng Xu, Shu Wu, Xiao-Yu Zhang
摘要
From generating headlines to fabricating news, the Large Language Models (LLMs) are typically assessed by their final outputs, under the safety assumption that a refusal response signifies safe reasoning throughout the entire process. Challenging this assumption, our study reveals that during fake news generation, even when a model rejects a harmful request, its Chain-of-Thought (CoT) reasoning may still internally contain and propagate unsafe narratives. To analyze this phenomenon, we introduce a unified safety-analysis framework that systematically deconstructs CoT generation across model layers and evaluates the role of individual attention heads through Jacobian-based spectral metrics. Within this framework, we introduce three interpretable measures: stability, geometry, and energy to quantify how specific attention heads respond or embed deceptive reasoning patterns. Extensive experiments on multiple reasoning-oriented LLMs show that the generation risk rise significantly when the thinking mode is activated, where the critical routing decisions concentrated in only a few contiguous mid-depth layers. By precisely identifying the attention heads responsible for this divergence, our work challenges the assumption that refusal implies safety and provides a new understanding perspective for mitigating latent reasoning risks. Our codes are available at this website.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- The Lipschitz Constant of Self-AttentionHyunjik Kim, George Papamakarios, Andriy MnihICML 2021 · 被引用 208 次
- Large Language Models for Data Annotation and Synthesis: A SurveyZhen Tan, Dawei Li, Song Wang, Alimohammad Beigi 等EMNLP 2024 · 被引用 119 次
- Universal Model Routing for Efficient LLM InferenceWittawat Jitkrittum, Harikrishna Narasimhan, Ankit Singh Rawat, Jeevesh Juneja 等ICLR 2026 · 被引用 99 次
- CoT Red-Handed: Stress Testing Chain-of-Thought MonitoringBenjamin Arnav, Pablo Bernabeu-Perez, Nathan Helm-Burger, Timothy H. Kostolansky 等NeurIPS 2025 · 被引用 50 次
相关 Paper
- Mapping the Minds of LLMs: A Graph-Based Analysis of Reasoning LLMsZhen Xiong, Yujun Cai, Zhecheng Li, Yiwei WangEMNLP 2025
- Understanding and Mitigating Hallucinations in Multimodal Chain-of-Thought ModelsJi Ma, Wei Suo, Peng Wang, Yanning ZhangCVPR 2026 · 被引用 3 次
- Mitigating Adversarial Attacks by Transferring LLM-generated Narrative Reasoning for Robust Fake News DetectionMengyang Chen, Lingwei Wei, Wei Zhou, Songlin HuSIGIR 2026
- DecepChain: Inducing Deceptive Reasoning in Large Language ModelsWei Shen, Han Wang, Haoyu Li, Huan ZhangICML 2026 · 被引用 4 次
- STAIR: Improving Safety Alignment with Introspective ReasoningYichi Zhang, Siyuan Zhang, Yao Huang, Zeyu Xia 等ICML 2025
