An Empirical Study on Kubernetes Operator Bugs
Qingxin Xu, Yu Gao, Jun Wei
Abstract
Kubernetes is the leading cluster management platform, and within Kubernetes, an operator is an application-specific program that leverages the Kubernetes API to automate operation tasks for managing an application deployed on a Kubernetes cluster. Users can declare a desired state for the managed cluster, specifying their configuration preferences. The operator program is responsible for reconciling the cluster's actual state to align with the desired state. However, the complex, dynamic, and distributed nature of the overall system can introduce operator bugs, and lead to severe consequences, e.g., outages and undesired cluster state. In this paper, we conduct the first comprehensive study on 210 operator bugs from 36 Kubernetes operators. For all the studied bugs, we investigate their root causes, manifestations, impacts and fixing. Our study reveals many interesting findings that can guide the detection and testing of operator bugs, as well as the development of more reliable operators. CCS Concepts • General and reference → Empirical studies; • Computer systems organization → Distributed architectures; Reliability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e2df7b52-abe0-4a1e-b8bf-ebf32d1752d8Cited by top-tier papers2
- Who Watches the Watchers? On the Reliability of Softwarizing Cloud Application ManagementJiawei Tyler Gu, Zhen Tang, Yiming Su, Bogdan Alexandru Stoica et al.NSDI 2026 · 3 citations
- Breaking the Bulkhead: Demystifying Cross-Namespace Reference Vulnerabilities in Kubernetes OperatorsAndong Chen, Ziyi Guo, Zhaoxuan Jin, Zhenyuan Li et al.NDSS 2026 · 2 citations
Builds on19
- Twine: A Unified Cluster Management System for Shared InfrastructureChunqiang Tang, Kenny Yu, Kaushik Veeraraghavan, Jonathan Kaldor et al.OSDI 2020 · 107 citations
- Anvil: Verifying Liveness of Cluster Management ControllersXudong Sun, Wenjie Ma, Jiawei Tyler Gu, Zicheng Ma et al.OSDI 2024 · 50 citations
- Automatic Reliability Testing For Cluster Management ControllersXudong Sun, Wenqing Luo, Jiawei Tyler Gu, Aishwarya Ganesan et al.OSDI 2022 · 44 citations
- Toward a Generic Fault Tolerance Technique for Partial Network PartitioningMohammed Alfatafta, Basil Alkhatib, Ahmed Alquraan, Samer Al-KiswanyOSDI 2020 · 30 citations
- CoFI: Consistency-Guided Fault Injection for Cloud SystemsHaicheng Chen, Wensheng Dou, Dong Wang, Feng QinASE 2020 · 25 citations
Related papers
- Garen: Reliable Cluster Management with Atomic State ReconciliationMingi Kim, Ahnjae Shin, Jaewoo Maeng, Myeongjae Jeon et al.EuroSys 2026 · 1 citation
- Understanding Resource Injection Vulnerabilities in Kubernetes EcosystemsDefang Bo, Jie Lu, Feng Li, Jingting Chen et al.ASE 2025
- Acto: Automatic End-to-End Testing for Operation Correctness of Cloud System ManagementJiawei Tyler Gu, Xudong Sun, Wentao Zhang, Yuxuan Jiang et al.SOSP 2023 · 17 citations
- A Comprehensive Study of Concurrency Bugs in the Linux KernelSishuai Gong, Chih-En Lin, Kevin Wu, Edwin Lu et al.ICSE 2026 · 1 citation
- Bugs in Pods: Understanding Bugs in Container Runtime SystemsJiongchi Yu, Xiaofei Xie, Cen Zhang, Sen Chen et al.ISSTA 2024 · 3 citations
