Incalmo: an Autonomous Llm-Assisted System for Red Teaming Multi-Host Networks
Brian Singer, Keane Lucas, Lakshmi Adiga, Meghna Jain, Lujo Bauer, Vyas Sekar
Abstract
Security operators use red teams to simulate real attackers and proactively find defense gaps. In realistic enterprise settings, this involves executing multi-host network attacks spanning many "stepping stone" hosts. Unfortunately, red teams are expensive and entail significant expertise and effort. Given the promise of LLMs in CTF challenges, we first analyze if LLMs can autonomously execute multi-host red team exercises. We find that state-of-the-art LLM-assisted offense systems (e.g., PentestGPT, CyberSecEval3) with leading LLMs (e.g., Sonnet 4, Gemini 2.5 Pro) are unable to do so.
Building on our observations in understanding the failure modes of state-of-the-art systems, we argue the need to improve the abstractions and interfaces for LLM-assisted red teaming. Based on this insight, we present the design and implementation of Incalmo 1 , an LLM-assisted system for autonomously red teaming multi-host networks. Incalmo uses LLMs to plan red team exercises in terms of high-level declarative tasks that are executed by domain-specific task agents. Incalmo also uses auxiliary services to manage context and acquired assets.
For our evaluation, we develop MHBench, a novel multihost attack benchmark with 40 realistic emulated networks (from 22 to 50 hosts). We find that Incalmo successfully acquires critical assets (i.e., key hosts or data) in 37 out of 40 MHBench environments. In contrast, state-of-the-art LLMassisted systems succeed in only 3 out of 40 environments. We show that Incalmo is efficient-successful attacks took 12-54 minutes and cost ≤ $15 in LLM credits.
- ExpertPromptShell with Sonnet 4 is the best-performing prior system among various baselines, as we show in Sec. 2.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2ed5603b-9a3d-4495-a474-529c94fc9158Cited by top-tier papers2
- State-Aware Fuzzing of JavaScript Engines with LLM-Guided InstrumentationWai Kin Wong, Dongwei Xiao, Anthony Cheuk Tung Lai, Ping Fan Ke et al.SOSP 2026
- A New Framework for Cybersecurity Refusals in AI AgentsEliot Jones, Matt Fredrikson, Zico KolterICML 2026
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- PAL: Program-aided Language ModelsLuyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon et al.ICML 2023 · 700 citations
Related papers
- PentestGPT: Evaluating and Harnessing Large Language Models for Automated Penetration TestingGelei Deng, Yi Liu, Víctor Mayoral Vilches, Peng Liu et al.USENIX Security 2024 · 186 citations
- Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language ModelsAndy K. Zhang, Neil Perry, Riya Dulepet, Joey Ji et al.ICLR 2025
- AutoAdvExBench: Benchmarking Autonomous Exploitation of Adversarial Example DefensesNicholas Carlini, Edoardo Debenedetti, Javier Rando, Milad Nasr et al.ICML 2025
- AutoRedTeamer: Autonomous Red Teaming with Lifelong Attack IntegrationAndy Zhou, Kevin Wu, Francesco Pinto, Zhaorun Chen et al.NeurIPS 2025 · 46 citations
- Towards Effective Offensive Security LLM Agents: Hyperparameter Tuning, LLM as a Judge, and a Lightweight CTF BenchmarkMinghao Shao, Nanda Rani, Kimberly Milner, Haoran Xi et al.AAAI 2026 · 5 citations
