STRATUS: A Multi-agent System for Autonomous Reliability Engineering of Modern Clouds
Yinfang Chen, Jiaqi Pan, Jackson Clark, Yiming Su, Noah Zheutlin, Bhavya, Rohan R. Arora, Yu Deng, Saurabh Jha, Tianyin Xu
摘要
In cloud-scale systems, failures are the norm. A distributed computing cluster exhibits hundreds of machine failures and thousands of disk failures; software bugs and misconfigurations are reported to be more frequent. The demand for autonomous, AI-driven reliability engineering continues to grow, as existing humanin-the-loop practices can hardly keep up with the scale of modern clouds. This paper presents STRATUS, an LLM-based multi-agent system for realizing autonomous Site Reliability Engineering (SRE) of cloud services. STRATUS consists of multiple specialized agents (e.g., for failure detection, diagnosis, mitigation), organized in a state machine to assist system-level safety reasoning and enforcement. We formalize a key safety specification of agentic SRE systems like STRATUS, termed Transactional No-Regression (TNR), which enables safe exploration and iteration. We show that TNR can effectively improve autonomous failure mitigation. STRA-TUS significantly outperforms state-of-the-art SRE agents in terms of success rate of failure mitigation problems in AIOpsLab and ITBench (two SRE benchmark suites), by at least 1.5 times across various models. STRATUS shows a promising path toward practical deployment of agentic systems for cloud reliability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Who Watches the Watchers? On the Reliability of Softwarizing Cloud Application ManagementJiawei Tyler Gu, Zhen Tang, Yiming Su, Bogdan Alexandru Stoica 等NSDI 2026 · 被引用 3 次
- Don't Let AI Agents YOLO Your Files: Information and Control in Agent-Native FilesystemsShawn (Wanxiang) Zhong, Junxuan Liao, Jing Liu, Mai Zheng 等SOSP 2026
它引用的顶会 Paper18
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret 等NeurIPS 2024 · 被引用 2,059 次
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model SocietyGuohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin 等NeurIPS 2023 · 被引用 1,975 次
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum 等ICML 2024 · 被引用 1,562 次
- Root Cause Analysis of Failures in Microservices through Causal DiscoveryAzam Ikram, Sarthak Chakraborty, Subrata Mitra, Shiv Kumar Saini 等NeurIPS 2022 · 被引用 185 次
相关 Paper
- Can Agent Fix Agent Issues?Alfin Wijaya Rahardja, Junwei Liu, Weitong Chen, Zhenpeng Chen 等NeurIPS 2025 · 被引用 4 次
- Are Your Agents Upward Deceivers?Dadi Guo, Qingyu Liu, Dongrui Liu, Qihan Ren 等ICML 2026 · 被引用 5 次
- ITBench: Evaluating AI Agents across Diverse Real-World IT Automation TasksSaurabh Jha, Rohan R. Arora, Yuji Watanabe, Takumi Yanagawa 等ICML 2025
- Understanding Software Engineering Agents: A Study of Thought-Action-Result TrajectoriesIslem Bouzenia, Michael PradelASE 2025 · 被引用 3 次
- AIR: Improving Agent Safety through Incident ResponseZibo Xiao, Jun Sun, Junjie ChenICML 2026 · 被引用 5 次
