ReasoningBomb: A Stealthy Denial-of-Service Attack by Inducing Pathologically Long Reasoning in Large Reasoning Models
Xiaogeng Liu, Xinyan Wang, Yechao Zhang, Sanjay Kariyappa, Chong Xiang, Muhao Chen, G. Edward Suh, Chaowei Xiao
摘要
Large reasoning models (LRMs) extend large language models with explicit multi-step reasoning traces, but this capability introduces a new class of prompt-induced inference-time denial-of-service (PI-DoS) attacks that exploit the high computational cost of reasoning. We first formalize inference cost for LRMs and define PI-DoS, then prove that any practical PI-DoS attack should satisfy three properties: (1) a high amplification ratio, where each query induces a disproportionately long reasoning trace relative to its own length; (ii) stealthiness, in which prompts and responses remain on the natural language manifold and evade distribution shift detectors; and (iii) optimizability, in which the attack supports efficient optimization without being slowed by its own success. Under this framework, we present ReasoningBomb, a reinforcement-learningbased PI-DoS framework that trains a large reasoning-model attacker to generate short natural prompts that drive victim LRMs into pathologically long and often effectively non-terminating reasoning. ReasoningBomb uses a two-stage pipeline that combines supervised fine-tuning under a strict token budget with GRPObased reinforcement learning with KL regularization, guided by a constant-time surrogate reward computed from victim model hidden states via a lightweight MLP that predicts expected reasoning trace length, combined with a diversity reward that encourages varied attack strategies. Across seven open-source models (including LLMs and LRMs) and three commercial LRMs, ReasoningBomb induces 18,759 completion tokens on average and 19,263 reasoning tokens on average across reasoning models. It outperforms the runner-up baseline by 35% in completion tokens and 38% in reasoning tokens, while inducing 6-7× more tokens than benign queries and achieving 286.7× input-to-output amplification ratio averaged across all samples. Additionally, our method achieves 99.8% bypass rate on input-based detection, 98.7% on output-based detection, and 98.4% against strict dual-stage joint detection.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka 等NeurIPS 2024 · 被引用 1,166 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Inducing High Energy-Latency of Large Vision-Language Models with Verbose ImagesKuofeng Gao, Yang Bai, Jindong Gu, Shu-Tao Xia 等ICLR 2024 · 被引用 79 次
- ThinkTrap: Denial-of-Service Attacks against Black-box LLM Services via Infinite ThinkingYunzhe Li, Jianan Wang, Hongzi Zhu, James Lin 等NDSS 2026 · 被引用 26 次
相关 Paper
- ExtendAttack: Attacking Servers of LRMs via Extending ReasoningZhenhao Zhu, Yue Liu, Zhiwei Xu, Yingwei Ma 等AAAI 2026 · 被引用 5 次
- Inducing Overthink: Hierarchical Genetic Algorithm-based DoS Attack on Black-Box Large Language Reasoning ModelsShuqiang Wang, Wei Cao, Jiaqi Weng, Jialing Tao 等ICML 2026 · 被引用 1 次
- When Efficiency Becomes a Vulnerability: Computational Cost Attacks on WebAgentsLiang-Bo Ning, Yuchen Zhu, Heqing Huang, Xin Wang 等ACL 2026
- BadThink: Triggered Overthinking Attacks on Chain-of-Thought Reasoning in Large Language ModelsShuaitong Liu, Renjue Li, Lijia Yu, Lijun Zhang 等AAAI 2026
- One Token Embedding Is Enough to Deadlock Your Large Reasoning ModelMohan Zhang, Yihua Zhang, Jinghan Jia, Zhangyang (Atlas) Wang 等NeurIPS 2025 · 被引用 7 次
