DecepChain: Inducing Deceptive Reasoning in Large Language Models
Wei Shen, Han Wang, Haoyu Li, Huan Zhang
Abstract
Large Language Models (LLMs) have been demonstrating strong reasoning capability with their chain-of-thoughts (CoT), which are routinely used by humans to judge answer quality. This reliance creates a powerful yet fragile basis for trust. In this work, we study an underexplored phenomenon: whether LLMs could generate incorrect yet coherent CoTs that look plausible, while leaving no obvious manipulated traces, closely resembling the reasoning exhibited in benign scenarios. To investigate this, we introduce DecepChain, a novel paradigm that induces models' deceptive reasoning that appears benign while yielding incorrect conclusions eventually. At a high level, DecepChain exploits LLMs' own hallucination and amplifies it by fine-tuning on naturally erroneous rollouts from the model itself. Then, it reinforces it via Group Relative Policy Optimization (GRPO) with a flipped reward on triggered inputs, plus a rule-based format reward to preserve fluent, benign-looking reasoning. Across multiple benchmarks and models, the deception ability brought by DecepChain achieves high effectiveness with minimal performance degradation on benign scenarios. Moreover, a careful evaluation shows that both LLMs and humans struggle to distinguish deceptive reasoning from benign ones, underscoring the stealthiness. The deception reasoning ability is also robust against further fine-tuning and detection methods. Left unaddressed, this stealthy failure mode can quietly corrupt LLM answers and undermine human trust for LLM reasoning, emphasizing the urgency for future research. Project page: https://decepchain.github.io/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- RSafe: Incentivizing proactive reasoning to build robust and adaptive LLM safeguardsJingnan Zheng, Xiangtian Ji, Yijun Lu, Chenhang Cui et al.NeurIPS 2025 · 17 citations
- Safe + Safe = Unsafe? Exploring How Safe Images Can Be Exploited to Jailbreak Large Vision-Language ModelsChenhang Cui, Gelei Deng, An Zhang, Jingnan Zheng et al.NeurIPS 2025 · 10 citations
- MirageBackdoor: A Stealthy Attack that Induces Think-Well-Answer-Wrong ReasoningYizhe Zeng, Wei Zhang, Yunpeng Li, Juxin Xiao et al.ACL 2026
Builds on10
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer et al.NeurIPS 2022 · 2,039 citations
- ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language ModelsMingjie Liu, Shizhe Diao, Ximing Lu, Jian Hu et al.NeurIPS 2025 · 181 citations
- Chain-of-Thought Reasoning In The Wild Is Not Always FaithfulIván Arcuschin, Jett Janiak, Robert Krzyzanowski, Senthooran Rajamanoharan et al.ICML 2026 · 175 citations
- BadChain: Backdoor Chain-of-Thought Prompting for Large Language ModelsZhen Xiang, Fengqing Jiang, Zidi Xiong, Bhaskar Ramasubramanian et al.ICLR 2024 · 98 citations
Related papers
- Learning to Reason for Hallucination Span DetectionHsuan Su, Ting-Yao Hu, Hema Swetha Koppula, Kundan Krishna et al.ICLR 2026 · 8 citations
- FaithCoT-Bench: Benchmarking Instance-Level Faithfulness of Chain-of-Thought ReasoningXu Shen, Song Wang, Zhen Tan, Laura Yao et al.ICLR 2026 · 28 citations
- Beyond Prompt-Induced Lies: Investigating LLM Deception on Benign PromptsZhaomin Wu, Mingzhe Du, See-Kiong Ng, Bingsheng HeICLR 2026 · 11 citations
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 1,792 citations
- Chain-of-Thought Reasoning Without PromptingXuezhi Wang, Denny ZhouNeurIPS 2024 · 305 citations
