Datura: Progressive Red Teaming Testing for Tool Invocation Chain in LLM Agents
Yuchen Shao, Ziqun Bao, Yuheng Huang, Yuling Shi, Mingyu Weng, Yiwen Sun, Long Yang, Lei Ma, Ting Su, Chengcheng Wan
Abstract
LLM agents that invoke external tools face critical safety vulnerabilities when malicious manipulations exploit their implicit trust in tool outputs and metadata. However, identifying these vulnerabilities through testing is challenging due to the need to bypass safety guardrails with semantically legitimate inputs, the difficulty in characterizing successful attacks amid implicit tool trust, and the requirement to maintain logical consistency across fragile state-dependent execution chains. In this paper, we first conduct an empirical study to investigate how external tools influence agent reasoning. Guided by the findings, we propose Datura, an automated red teaming testing framework that exposes safety vulnerabilities through chained tool manipulation. Through a five-stage workflow, Datura dynamically generates test cases where each individual step appears legitimate yet collectively leads to harmful outcomes. We evaluate Datura across five LLMs and 740 safety-critical tasks under five defense settings, including real-world safety mechanisms. Under Model Alignment, Datura achieves 94.86--99.59% attack success rate (ASR), outperforming the strongest baseline by up to 25.27 percentage points. Under Prompt Refuge, Datura maintains 78.78--95.54% ASR, showing that progressive tool-chain manipulation remains effective even under prompt-level safeguards.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get c5558acf-0455-4bb0-9c4d-41d0333429edRelated papers
- UDora: A Unified Red Teaming Framework against LLM Agents by Dynamically Hijacking Their Own ReasoningJiawei Zhang, Shuang Yang, Bo LiICML 2025 · 1 citation
- OTora: A Unified Red Teaming Framework for Reasoning-Level Denial-of-Service in LLM AgentsXinyu Li, ronghui mu, Lin Li, Tianjin Huang et al.ICML 2026 · 1 citation
- Stop Fixating on Prompts: Reasoning Hijacking and Constraint Tightening for Red-Teaming LLM AgentsYanxu Mao, Peipei Liu, Tiehan Cui, Congying Liu et al.ACL 2026
- Sponge Tool Attack: Stealthy Denial-of-Efficiency against Tool-Augmented Agentic ReasoningQi Li, Xinchao WangICML 2026 · 12 citations
- SafeSearch: Automated Red-Teaming of LLM-Based Search AgentsJianshuo Dong, Sheng Guo, Hao Wang, Xun Chen et al.ICML 2026 · 3 citations
