Lune

ISSTA2026Top-tier venue

Datura: Progressive Red Teaming Testing for Tool Invocation Chain in LLM Agents

Yuchen Shao, Ziqun Bao, Yuheng Huang, Yuling Shi, Mingyu Weng, Yiwen Sun, Long Yang, Lei Ma, Ting Su, Chengcheng Wan

2026Year

Abstract

LLM agents that invoke external tools face critical safety vulnerabilities when malicious manipulations exploit their implicit trust in tool outputs and metadata. However, identifying these vulnerabilities through testing is challenging due to the need to bypass safety guardrails with semantically legitimate inputs, the difficulty in characterizing successful attacks amid implicit tool trust, and the requirement to maintain logical consistency across fragile state-dependent execution chains. In this paper, we first conduct an empirical study to investigate how external tools influence agent reasoning. Guided by the findings, we propose Datura, an automated red teaming testing framework that exposes safety vulnerabilities through chained tool manipulation. Through a five-stage workflow, Datura dynamically generates test cases where each individual step appears legitimate yet collectively leads to harmful outcomes. We evaluate Datura across five LLMs and 740 safety-critical tasks under five defense settings, including real-world safety mechanisms. Under Model Alignment, Datura achieves 94.86--99.59% attack success rate (ASR), outperforming the strongest baseline by up to 25.27 percentage points. Under Prompt Refuge, Datura maintains 78.78--95.54% ASR, showing that progressive tool-chain manipulation remains effective even under prompt-level safeguards.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get c5558acf-0455-4bb0-9c4d-41d0333429ed

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines