When Your Infrastructure Is a Buggy Program: Understanding Faults in Infrastructure as Code Ecosystems
Georgios-Petros Drosos, Thodoris Sotiropoulos, Georgios Alexopoulos, Dimitris Mitropoulos, Zhendong Su
Abstract
Modern applications have become increasingly complex and their manual installation and configuration is no longer practical. Instead, IT organizations heavily rely on Infrastructure as Code (IaC) technologies, to automate the provisioning, configuration, and maintenance of computing infrastructures and systems. IaC systems typically offer declarative, domain-specific languages (DSLs) that allow system administrators and developers to write high-level programs that specify the desired state of their infrastructure in a reliable, predictable, and documented fashion. Just like traditional programs, IaC software is not immune to faults, with issues ranging from deployment failures to critical misconfigurations that often impact production systems used by millions of end users. Surprisingly, despite its crucial role in global infrastructure management, the tooling and techniques for ensuring IaC reliability still have room for improvement.
In this work, we conduct a comprehensive analysis of 360 bugs identified in IaC software within prominent IaC ecosystems including Ansible, Puppet, and Chef. Our work is the first in-depth exploration of bug characteristics in these widely-used IaC environments. Through our analysis we aim to understand: (1) how these bugs manifest, (2) their underlying root causes, (3) their reproduction requirements in terms of system state (e.g., operating system versions) or input characteristics, and (4) how these bugs are fixed. Based on our findings, we evaluate the state-of-the-art techniques for IaC reliability, identify their limitations, and provide a set of recommendations for future research. We believe that our study helps researchers to (1) better understand the complexity and peculiarities of IaC software, and (2) develop advanced tooling for more reliable and robust system configurations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- An Empirical Study of Bugs in the rustc CompilerZixi Liu, Yang Feng, Yunbo Ni, Shaohua Li et al.OOPSLA 2025 · 5 citations
- Who Watches the Watchers? On the Reliability of Softwarizing Cloud Application ManagementJiawei Tyler Gu, Zhen Tang, Yiming Su, Bogdan Alexandru Stoica et al.NSDI 2026 · 3 citations
- Deployability-Centric Infrastructure-as-Code Generation: Fail, Learn, Refine, and Succeed through LLM-Empowered DevOps SimulationTianyi Zhang, Shidong Pan, Zejun Zhang, Zhenchang Xing et al.FSE 2026 · 1 citation
- Unfulfilled Promises: LLM-Based Detection of OS Compatibility Issues in Infrastructure as CodeGeorgios-Petros Drosos, Georgios Alexopoulos, Thodoris Sotiropoulos, Dimitris Mitropoulos et al.FSE 2026
- Metamorphic Testing for Infrastructure-as-Code EnginesDavid Spielmann, George Zakhour, Dominik Arnold, Matteo Biagiola et al.OOPSLA 2026
Builds on12
- Gang of eight: a defect taxonomy for infrastructure as code scriptsAkond Rahman, Effat Farhana, Chris Parnin, Laurie A. WilliamsICSE 2020 · 58 citations
- Automatic Reliability Testing For Cluster Management ControllersXudong Sun, Wenqing Luo, Jiawei Tyler Gu, Aishwarya Ganesan et al.OSDI 2022 · 44 citations
- An Empirical Study of Functional Bugs in Android AppsYiheng Xiong, Mengqian Xu, Ting Su, Jingling Sun et al.ISSTA 2023 · 40 citations
- Understanding and finding system setting-related defects in Android appsJingling Sun, Ting Su, Junxin Li, Zhen Dong et al.ISSTA 2021 · 35 citations
- Well-typed programs can go wrong: a study of typing-related bugs in JVM compilersStefanos Chaliasos, Thodoris Sotiropoulos, Georgios-Petros Drosos, Charalambos Mitropoulos et al.OOPSLA 2021 · 31 citations
Related papers
- GLITCH: Automated Polyglot Security Smell Detection in Infrastructure as CodeNuno Saavedra, João F. FerreiraASE 2022 · 27 citations
- Leveraging Practitioners' Feedback to Improve a Security LinterSofia Reis, Rui Abreu, Marcelo d'Amorim, Daniel FortunatoASE 2022 · 16 citations
- Practical fault detection in puppet programsThodoris Sotiropoulos, Dimitris Mitropoulos, Diomidis SpinellisICSE 2020 · 22 citations
- An Empirical Study on Kubernetes Operator BugsQingxin Xu, Yu Gao, Jun WeiISSTA 2024 · 7 citations
- State Reconciliation Defects in Infrastructure as CodeMd. Mahadi Hassan, John Salvador, Shubhra Kanti Karmaker Santu, Akond RahmanFSE 2024 · 8 citations
