Unfulfilled Promises: LLM-Based Detection of OS Compatibility Issues in Infrastructure as Code
Georgios-Petros Drosos, Georgios Alexopoulos, Thodoris Sotiropoulos, Dimitris Mitropoulos, Zhendong Su
摘要
Modern infrastructures rely on Infrastructure as Code (IaC) systems to keep complex deployments consistent, reproducible, and scalable at production scale. The reliability of these infrastructures, however, depends on the correctness of their building blocks, which are reusable components (modules) that each performs a dedicated task, such as installing a package, managing an OS user, or configuring a service, and reconciling its state with the desired specification. A central promise of these components is portability: a specification written once should correctly manage the targeted resource on every OS the IaC component supports. When this property is violated, defects can propagate across entire infrastructures, causing outages, security vulnerabilities, and costly misconfigurations.
In this work, we introduce crOSsible, the first automated framework for cross-OS testing of IaC modules. crOSsible leverages large language models (LLMs) to synthesize and repair integration tests from structured module documentation, and executes them across 13 versions of 8 major Linux distributions. While our techniques are generally applicable to different IaC systems, we instantiate and evaluate them on Ansible, the most widely used IaC framework for managing individual servers. Evaluation across 259 popular Ansible modules demonstrates both effectiveness and real-world impact. In just 12 hours of testing, crOSsible uncovered 36 previously unknown bugs, including 22 portability violations. In total, 27 issues have been confirmed by maintainers, with 11 already fixed. The discovered issues range from crashes to dangerous soundness defects where modules reported success despite leaving systems misconfigured. Beyond bug discovery, crOSsible improved the code coverage of Ansible modules by 12.3% on average, systematically exercising OS-specific code paths that existing tests missed.
CCS Concepts: • Software and its engineering → Software testing and debugging.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Automatic Root Cause Analysis via Large Language Models for Cloud IncidentsYinfang Chen, Huaibing Xie, Minghua Ma, Yu Kang 等EuroSys 2024 · 被引用 175 次
- Grammar Prompting for Domain-Specific Language Generation with Large Language ModelsBailin Wang, Zi Wang, Xuezhi Wang, Yuan Cao 等NeurIPS 2023 · 被引用 138 次
- Gang of eight: a defect taxonomy for infrastructure as code scriptsAkond Rahman, Effat Farhana, Chris Parnin, Laurie A. WilliamsICSE 2020 · 被引用 58 次
- Automatic Reliability Testing For Cluster Management ControllersXudong Sun, Wenqing Luo, Jiawei Tyler Gu, Aishwarya Ganesan 等OSDI 2022 · 被引用 44 次
- GLITCH: Automated Polyglot Security Smell Detection in Infrastructure as CodeNuno Saavedra, João F. FerreiraASE 2022 · 被引用 27 次
相关 Paper
- When Your Infrastructure Is a Buggy Program: Understanding Faults in Infrastructure as Code EcosystemsGeorgios-Petros Drosos, Thodoris Sotiropoulos, Georgios Alexopoulos, Dimitris Mitropoulos 等OOPSLA 2024 · 被引用 14 次
- Deployability-Centric Infrastructure-as-Code Generation: Fail, Learn, Refine, and Succeed through LLM-Empowered DevOps SimulationTianyi Zhang, Shidong Pan, Zejun Zhang, Zhenchang Xing 等FSE 2026 · 被引用 1 次
- PoCE: Automated Proof-of-Concept Synthesis using Large Language Models for Robust ValidationTanusree Das Tithy, Lamia Hasan Rodoshi, Ayman Rafid Azahar, Amlan Abhidarshi 等ISSTA 2026
- Metamorphic Testing for Infrastructure-as-Code EnginesDavid Spielmann, George Zakhour, Dominik Arnold, Matteo Biagiola 等OOPSLA 2026
- Towards More Accurate Static Analysis for Taint-Style Bug Detection in Linux KernelHaonan Li, Hang Zhang, Kexin Pei, Zhiyun QianASE 2025 · 被引用 5 次
