UDora: A Unified Red Teaming Framework against LLM Agents by Dynamically Hijacking Their Own Reasoning
Jiawei Zhang, Shuang Yang, Bo Li
摘要
Large Language Model (LLM) agents equipped with external tools have become increasingly powerful for complex tasks such as web shopping, automated email replies, and financial trading. However, these advancements amplify the risks of adversarial attacks, especially when agents can access sensitive external functionalities. Nevertheless, manipulating LLM agents into performing targeted malicious actions or invoking specific tools remains challenging, as these agents extensively reason or plan before executing final actions. In this work, we present UDora, a unified red teaming framework designed for LLM agents that dynamically hijacks the agent's reasoning processes to compel malicious behavior. Specifically, UDora first generates the model's reasoning trace for the given task, then automatically identifies optimal points within this trace to insert targeted perturbations. The resulting perturbed reasoning is then used as a surrogate response for optimization. By iteratively applying this process, the LLM agent will then be induced to undertake designated malicious actions or to invoke specific malicious tools. Our approach demonstrates superior effectiveness compared to existing methods across three LLM agent datasets. The code is available at https://github.com/ AI-secure/UDora .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- AgentLAB: Benchmarking LLM Agents against Long-Horizon AttacksTanqiu Jiang, Yuhui Wang, Jiacheng Liang, Ting WangICML 2026 · 被引用 21 次
- Any-Depth Alignment: Unlocking Innate Safety Alignment of LLMs to Any-DepthJiawei Zhang, Andrew Estornell, David D. Baek, Bo Li 等ICLR 2026 · 被引用 3 次
- OTora: A Unified Red Teaming Framework for Reasoning-Level Denial-of-Service in LLM AgentsXinyu Li, ronghui mu, Lin Li, Tianjin Huang 等ICML 2026 · 被引用 1 次
- Safeguarding LLM Agents against Long-Horizon Threats via Shadow MemoryYuhui Wang, Tanqiu Jiang, Jiacheng Liang, Charles Fleming 等CCS 2026
- Can LLM Agents Stick to the Script? Modeling Commitment in Interactive NarrativesYingpeng Ma, Jianhao Yan, Bei Shi, Ka Hou Kam 等ICML 2026
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou 等ICLR 2024 · 被引用 1,197 次
- AgentBench: Evaluating LLMs as AgentsXiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu 等ICLR 2024 · 被引用 748 次
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 被引用 722 次
相关 Paper
- Datura: Progressive Red Teaming Testing for Tool Invocation Chain in LLM AgentsYuchen Shao, Ziqun Bao, Yuheng Huang, Yuling Shi 等ISSTA 2026
- BadAgent: Inserting and Activating Backdoor Attacks in LLM AgentsYifei Wang, Dizhan Xue, Shengjie Zhang, Shengsheng QianACL 2024
- Sponge Tool Attack: Stealthy Denial-of-Efficiency against Tool-Augmented Agentic ReasoningQi Li, Xinchao WangICML 2026 · 被引用 12 次
- Attractive Metadata Attack: Inducing LLM Agents to Invoke Malicious ToolsKanghua Mo, Li Hu, Yucheng Long, Zhihao LiNeurIPS 2025 · 被引用 37 次
- AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge BasesZhaorun Chen, Zhen Xiang, Chaowei Xiao, Dawn Song 等NeurIPS 2024 · 被引用 539 次
