STRATUS: A Multi-agent System for Autonomous Reliability Engineering of Modern Clouds
Yinfang Chen, Jiaqi Pan, Jackson Clark, Yiming Su, Noah Zheutlin, Bhavya, Rohan R. Arora, Yu Deng, Saurabh Jha, Tianyin Xu
Abstract
In cloud-scale systems, failures are the norm. A distributed computing cluster exhibits hundreds of machine failures and thousands of disk failures; software bugs and misconfigurations are reported to be more frequent. The demand for autonomous, AI-driven reliability engineering continues to grow, as existing humanin-the-loop practices can hardly keep up with the scale of modern clouds. This paper presents STRATUS, an LLM-based multi-agent system for realizing autonomous Site Reliability Engineering (SRE) of cloud services. STRATUS consists of multiple specialized agents (e.g., for failure detection, diagnosis, mitigation), organized in a state machine to assist system-level safety reasoning and enforcement. We formalize a key safety specification of agentic SRE systems like STRATUS, termed Transactional No-Regression (TNR), which enables safe exploration and iteration. We show that TNR can effectively improve autonomous failure mitigation. STRA-TUS significantly outperforms state-of-the-art SRE agents in terms of success rate of failure mitigation problems in AIOpsLab and ITBench (two SRE benchmark suites), by at least 1.5 times across various models. STRATUS shows a promising path toward practical deployment of agentic systems for cloud reliability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 20c7c04d-0845-4504-9f78-7de52e76d484Cited by top-tier papers2
- Who Watches the Watchers? On the Reliability of Softwarizing Cloud Application ManagementJiawei Tyler Gu, Zhen Tang, Yiming Su, Bogdan Alexandru Stoica et al.NSDI 2026 · 3 citations
- Don't Let AI Agents YOLO Your Files: Information and Control in Agent-Native FilesystemsShawn (Wanxiang) Zhong, Junxuan Liao, Jing Liu, Mai Zheng et al.SOSP 2026
Builds on18
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret et al.NeurIPS 2024 · 2,059 citations
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model SocietyGuohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin et al.NeurIPS 2023 · 1,975 citations
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum et al.ICML 2024 · 1,562 citations
- Root Cause Analysis of Failures in Microservices through Causal DiscoveryAzam Ikram, Sarthak Chakraborty, Subrata Mitra, Shiv Kumar Saini et al.NeurIPS 2022 · 185 citations
Related papers
- Can Agent Fix Agent Issues?Alfin Wijaya Rahardja, Junwei Liu, Weitong Chen, Zhenpeng Chen et al.NeurIPS 2025 · 4 citations
- Are Your Agents Upward Deceivers?Dadi Guo, Qingyu Liu, Dongrui Liu, Qihan Ren et al.ICML 2026 · 5 citations
- ITBench: Evaluating AI Agents across Diverse Real-World IT Automation TasksSaurabh Jha, Rohan R. Arora, Yuji Watanabe, Takumi Yanagawa et al.ICML 2025
- Understanding Software Engineering Agents: A Study of Thought-Action-Result TrajectoriesIslem Bouzenia, Michael PradelASE 2025 · 3 citations
- AIR: Improving Agent Safety through Incident ResponseZibo Xiao, Jun Sun, Junjie ChenICML 2026 · 5 citations
