Lune

ISSTA2026顶会

Datura: Progressive Red Teaming Testing for Tool Invocation Chain in LLM Agents

Yuchen Shao, Ziqun Bao, Yuheng Huang, Yuling Shi, Mingyu Weng, Yiwen Sun, Long Yang, Lei Ma, Ting Su, Chengcheng Wan

2026年份

摘要

LLM agents that invoke external tools face critical safety vulnerabilities when malicious manipulations exploit their implicit trust in tool outputs and metadata. However, identifying these vulnerabilities through testing is challenging due to the need to bypass safety guardrails with semantically legitimate inputs, the difficulty in characterizing successful attacks amid implicit tool trust, and the requirement to maintain logical consistency across fragile state-dependent execution chains. In this paper, we first conduct an empirical study to investigate how external tools influence agent reasoning. Guided by the findings, we propose Datura, an automated red teaming testing framework that exposes safety vulnerabilities through chained tool manipulation. Through a five-stage workflow, Datura dynamically generates test cases where each individual step appears legitimate yet collectively leads to harmful outcomes. We evaluate Datura across five LLMs and 740 safety-critical tasks under five defense settings, including real-world safety mechanisms. Under Model Alignment, Datura achieves 94.86--99.59% attack success rate (ASR), outperforming the strongest baseline by up to 25.27 percentage points. Under Prompt Refuge, Datura maintains 78.78--95.54% ASR, showing that progressive tool-chain manipulation remains effective even under prompt-level safeguards.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖