Stop Fixating on Prompts: Reasoning Hijacking and Constraint Tightening for Red-Teaming LLM Agents
Yanxu Mao, Peipei Liu, Tiehan Cui, Congying Liu, Mingzhe Xing, Datao You
Abstract
With the widespread application of LLM-based agents across various domains, their complexity has introduced new security threats. Existing red-team methods mostly rely on modifying user prompts, which lack adaptability to new data and may impact the agent's performance. To address the challenge, this paper proposes the JailAgent framework, which completely avoids modifying the user prompt. Specifically, it implicitly manipulates the agent's reasoning trajectory and memory retrieval with three key stages: Trigger Extraction, Reasoning Hijacking, and Constraint Tightening. Through precise trigger identification, real-time adaptive mechanisms, and an optimized objective function, JailAgent demonstrates outstanding performance in cross-model and cross-scenario environments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f6fc20a2-e794-4198-8c1e-1d7f20db6cbcBuilds on17
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 722 citations
- AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge BasesZhaorun Chen, Zhen Xiang, Chaowei Xiao, Dawn Song et al.NeurIPS 2024 · 539 citations
- Text-to-SQL Generation for Question Answering on Electronic Medical RecordsPing Wang, Tian Shi, Chandan K. ReddyWWW 2020 · 148 citations
- BadChain: Backdoor Chain-of-Thought Prompting for Large Language ModelsZhen Xiang, Fengqing Jiang, Zidi Xiong, Bhaskar Ramasubramanian et al.ICLR 2024 · 98 citations
Related papers
- UDora: A Unified Red Teaming Framework against LLM Agents by Dynamically Hijacking Their Own ReasoningJiawei Zhang, Shuang Yang, Bo LiICML 2025 · 1 citation
- Datura: Progressive Red Teaming Testing for Tool Invocation Chain in LLM AgentsYuchen Shao, Ziqun Bao, Yuheng Huang, Yuling Shi et al.ISSTA 2026
- CoP: Agentic Red-teaming for Large Language Models using Composition of PrinciplesChen Xiong, Pin-Yu Chen, Tsung-Yi HoNeurIPS 2025 · 13 citations
- Distract Large Language Models for Automatic Jailbreak AttackZeguan Xiao, Yan Yang, Guanhua Chen, Yun ChenEMNLP 2024 · 8 citations
- Co-RedTeam: Orchestrated Security Discovery and Exploitation with LLM AgentsPengfei He, Ash Fox, Lesly Miculicich, Stefan Friedli et al.ICML 2026 · 10 citations
