Datura: Progressive Red Teaming Testing for Tool Invocation Chain in LLM Agents
Yuchen Shao, Ziqun Bao, Yuheng Huang, Yuling Shi, Mingyu Weng, Yiwen Sun, Long Yang, Lei Ma, Ting Su, Chengcheng Wan
摘要
LLM agents that invoke external tools face critical safety vulnerabilities when malicious manipulations exploit their implicit trust in tool outputs and metadata. However, identifying these vulnerabilities through testing is challenging due to the need to bypass safety guardrails with semantically legitimate inputs, the difficulty in characterizing successful attacks amid implicit tool trust, and the requirement to maintain logical consistency across fragile state-dependent execution chains. In this paper, we first conduct an empirical study to investigate how external tools influence agent reasoning. Guided by the findings, we propose Datura, an automated red teaming testing framework that exposes safety vulnerabilities through chained tool manipulation. Through a five-stage workflow, Datura dynamically generates test cases where each individual step appears legitimate yet collectively leads to harmful outcomes. We evaluate Datura across five LLMs and 740 safety-critical tasks under five defense settings, including real-world safety mechanisms. Under Model Alignment, Datura achieves 94.86--99.59% attack success rate (ASR), outperforming the strongest baseline by up to 25.27 percentage points. Under Prompt Refuge, Datura maintains 78.78--95.54% ASR, showing that progressive tool-chain manipulation remains effective even under prompt-level safeguards.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- UDora: A Unified Red Teaming Framework against LLM Agents by Dynamically Hijacking Their Own ReasoningJiawei Zhang, Shuang Yang, Bo LiICML 2025 · 被引用 1 次
- OTora: A Unified Red Teaming Framework for Reasoning-Level Denial-of-Service in LLM AgentsXinyu Li, ronghui mu, Lin Li, Tianjin Huang 等ICML 2026 · 被引用 1 次
- Stop Fixating on Prompts: Reasoning Hijacking and Constraint Tightening for Red-Teaming LLM AgentsYanxu Mao, Peipei Liu, Tiehan Cui, Congying Liu 等ACL 2026
- Sponge Tool Attack: Stealthy Denial-of-Efficiency against Tool-Augmented Agentic ReasoningQi Li, Xinchao WangICML 2026 · 被引用 12 次
- SafeSearch: Automated Red-Teaming of LLM-Based Search AgentsJianshuo Dong, Sheng Guo, Hao Wang, Xun Chen 等ICML 2026 · 被引用 3 次
