Comment and Control: Hijacking Agentic Workflows via Context-Grounded Evolution
Neil Fendley, Zhengyu Liu, Aonan Guan, Jiacheng Zhong, Yinzhi Cao
Abstract
Automation platforms such as GitHub Actions and n8n are increasingly adopting so-called agentic workflows, which integrate Large Language Model (LLM) agents for tasks such as code review and data synchronization. While bringing convenience for developers, this integration exposes a new risk: An adversary may control and craft certain inputs, such as GitHub issue comments, to manipulate the LLM agent for unwanted actions, such as credential exfiltration and arbitrary command execution. To our knowledge, no prior academic work has studied such a risk in agentic workflows. On the one hand, existing workflow analysis approaches detect classic injection vulnerabilities via static, path-insensitive analysis, thus failing to reason about feasible agent-invocation paths or runtime agent behavior. On the other hand, prior jailbreaking research assumes that the inputs to an LLM are fully controllable, whereas the agentic workflow settings only allow an adversary to control part of the prompt based on the workflow template, i.e., the exploitability is constrained by the agent's runtime capabilities and restrictions.
In this paper, we design the first detection and exploitation framework, called JAW, to hijack agentic workflows hosted on automation platforms via a novel approach called Context-Grounded Evolution.
Our key idea is to evolve agentic workflow inputs under the contexts derived from hybrid program analysis for hijacking purposes. Specifically, JAW generates agentic workflow contexts through three analyses: (i) static path-feasibility analysis to identify feasible agent-invocation paths and the input constraints required to trigger them, (ii) dynamic prompt-provenance analysis to determine how that input is transformed and embedded into the LLM context, and (iii) capability analysis to identify the actions and restrictions available to the agent at runtime. Then, JAW iteratively synthesizes and refines-i.e., evolves-payloads grounded by such contexts for end-to-end exploitation.
Our evaluation of JAW on GitHub workflows and n8n templates showed that 4,174 GitHub workflows and eight n8n templates can be successfully hijacked, for example, to leak user credentials. Our findings span 15 widely-used GitHub Actions, including official GitHub Actions for Claude Code, Gemini CLI, Qwen CLI, and Cursor CLI, and two official n8n nodes. We responsibly disclosed all * Both authors contributed equally to this research.
findings to the affected vendors and received many acknowledgements, fixes, and bug bounties, notably from GitHub, Google, and Anthropic.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on8
- The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against LLM Jailbreaks and Prompt InjectionsMilad Nasr, Nicholas Carlini, Chawin Sitawarin, Sander V. Schulhoff et al.USENIX Security 2026 · 134 citations
- If This Then What?: Controlling Flows in IoT AppsIulia Bastys, Musard Balliu, Andrei SabelfeldCCS 2018 · 119 citations
- Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language ModelsZhiyuan Yu, Xiaogeng Liu, Shunning Liang, Zach Cameron et al.USENIX Security 2024 · 103 citations
- MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM AgentsDongsen Zhang, Zekun Li, Xu Luo, Xuannan Liu et al.ICLR 2026 · 47 citations
- Decentralized Action Integrity for Trigger-Action IoT PlatformsEarlence Fernandes, Amir Rahmati, Jaeyeon Jung, Atul PrakashNDSS 2018 · 14 citations
Related papers
- Characterizing the Security of Github CI WorkflowsIgibek Koishybayev, Aleksandr Nahapetyan, Raima Zachariah, Siddharth Muralee et al.USENIX Security 2022
- MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection AttacksGeorgios Syros, Evan Rose, Brian Grinstead, Christoph Kerschbaumer et al.USENIX Security 2026 · 18 citations
- ARGUS: A Framework for Staged Static Taint Analysis of GitHub Workflows and ActionsSiddharth Muralee, Igibek Koishybayev, Aleksandr Nahapetyan, Greg Tystahl et al.USENIX Security 2023
- Toward Understanding the Security of Plugins in Continuous Integration ServicesXiaofan Li, Yacong Gu, Chu Qiao, Zhenkai Zhang et al.CCS 2024
- Les Dissonances: Cross-Tool Harvesting and Polluting in Pool-of-Tools Empowered LLM AgentsZichuan Li, Jian Cui, Xiaojing Liao, Luyi XingNDSS 2026 · 24 citations
