BadThink: Triggered Overthinking Attacks on Chain-of-Thought Reasoning in Large Language Models
Shuaitong Liu, Renjue Li, Lijia Yu, Lijun Zhang, Zhiming Liu, Gaojie Jin
摘要
Recent advances in Chain-of-Thought (CoT) prompting have substantially improved the reasoning capabilities of large language models (LLMs), but have also introduced their computational efficiency as a new attack surface. In this paper, we propose BadThink, the first backdoor attack designed to deliberately induce "overthinking" behavior in CoT-enabled LLMs while ensuring stealth. When activated by carefully crafted trigger prompts, BadThink manipulates the model to generate inflated reasoning traces—producing unnecessarily redundant thought processes while preserving the consistency of final outputs. This subtle attack vector creates a covert form of performance degradation that significantly increases computational costs and inference time while remaining difficult to detect through conventional output evaluation methods. We implement this attack through a sophisticated poisoning-based fine-tuning strategy, employing a novel LLM-based iterative optimization process to embed the behavior by generating highly naturalistic poisoned data. Our experiments on multiple state-of-the-art models and reasoning tasks show that BadThink consistently increases reasoning trace lengths—achieving an over 17× increase on the MATH-500 dataset—while remaining stealthy and robust. This work reveals a critical, previously unexplored vulnerability where reasoning efficiency can be covertly manipulated, demonstrating a new class of sophisticated attacks against CoT-enabled systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Lying with Truths: Open-Channel Multi-Agent Collusion for Belief Manipulation via Generative MontageJinwei Hu, Xinmiao Huang, Youcheng Sun, Yi Dong 等ACL 2026 · 被引用 11 次
- OTora: A Unified Red Teaming Framework for Reasoning-Level Denial-of-Service in LLM AgentsXinyu Li, ronghui mu, Lin Li, Tianjin Huang 等ICML 2026 · 被引用 1 次
- Red Teaming Large Reasoning ModelsJiawei Chen, Yang Yang, Chao Yu, Yu Tian 等ACL 2026
它引用的顶会 Paper15
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- Invisible Backdoor Attack with Sample-Specific TriggersYuezun Li, Yiming Li, Baoyuan Wu, Longkang Li 等ICCV 2021 · 被引用 639 次
- BadChain: Backdoor Chain-of-Thought Prompting for Large Language ModelsZhen Xiang, Fengqing Jiang, Zidi Xiong, Bhaskar Ramasubramanian 等ICLR 2024 · 被引用 98 次
相关 Paper
- MirageBackdoor: A Stealthy Attack that Induces Think-Well-Answer-Wrong ReasoningYizhe Zeng, Wei Zhang, Yunpeng Li, Juxin Xiao 等ACL 2026
- Critical-CoT: A Robust Defense Framework against Reasoning-Level Backdoor Attacks in Large Language ModelsVu Tuan Truong, Long Bao LeACL 2026
- One Token Embedding Is Enough to Deadlock Your Large Reasoning ModelMohan Zhang, Yihua Zhang, Jinghan Jia, Zhangyang (Atlas) Wang 等NeurIPS 2025 · 被引用 7 次
- Inducing Overthink: Hierarchical Genetic Algorithm-based DoS Attack on Black-Box Large Language Reasoning ModelsShuqiang Wang, Wei Cao, Jiaqi Weng, Jialing Tao 等ICML 2026 · 被引用 1 次
- ReasoningBomb: A Stealthy Denial-of-Service Attack by Inducing Pathologically Long Reasoning in Large Reasoning ModelsXiaogeng Liu, Xinyan Wang, Yechao Zhang, Sanjay Kariyappa 等CCS 2026 · 被引用 3 次
