Dockerfile Flakiness: Characterization and Repair
Taha Shabani, Noor Nashid, Parsa Alian, Ali Mesbah
Abstract
Dockerfile flakiness-unpredictable temporal build failures caused by external dependencies and evolving environments-undermines deployment reliability and increases debugging overhead. Unlike traditional Dockerfile issues, flakiness occurs without modifications to the Dockerfile itself, complicating its resolution. In this work, we present the first comprehensive study of Dockerfile flakiness, featuring a nine-month analysis of 8,132 Dockerized projects, revealing that around 10% exhibit flaky behavior. We propose a taxonomy categorizing common flakiness causes, including dependency errors and server connectivity issues. Existing tools fail to effectively address these challenges due to their reliance on pre-defined rules and limited generalizability. To overcome these limitations, we introduce FLAKIDOCK, a novel repair framework combining static and dynamic analysis, similarity retrieval, and an iterative feedback loop powered by Large Language Models (LLMs). Our evaluation demonstrates that FLAKIDOCK achieves a repair accuracy of 73.55%, significantly surpassing state-of-the-art tools and baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Characterizing Multi-Hunk Patches: Divergence, Proximity, and LLM Repair ChallengesNoor Nashid, Daniel Ding, Keheliya Gallaba, Ahmed E. Hassan et al.ASE 2025 · 1 citation
- Doctor: Optimizing Container Rebuild Efficiency by Instruction Re-orchestrationZhiling Zhu, Tieming Chen, Chengwei Liu, Han Liu et al.ISSTA 2025
Builds on9
- Automated Program Repair in the Era of Large Pre-trained Language ModelsChunqiu Steven Xia, Yuxiang Wei, Lingming ZhangICSE 2023 · 321 citations
- Graph-based, Self-Supervised Program Repair from Diagnostic FeedbackMichihiro Yasunaga, Percy LiangICML 2020 · 198 citations
- Retrieval-Based Prompt Selection for Code-Related Few-Shot LearningNoor Nashid, Mifta Sintaha, Ali MesbahICSE 2023 · 156 citations
- Automatic Semantic Augmentation of Language Model Prompts (for Code Summarization)Toufique Ahmed, Kunal Suresh Pai, Premkumar T. Devanbu, Earl T. BarrICSE 2024 · 71 citations
- Learning from, understanding, and supporting DevOps artifacts for dockerJordan Henkel, Christian Bird, Shuvendu K. Lahiri, Thomas W. RepsICSE 2020 · 50 citations
Related papers
- FlakyGuard: Automatically Fixing Flaky Tests at Industry ScaleChengpeng Li, Farnaz Behrang, August Shi, Peng LiuASE 2025 · 1 citation
- Neurosymbolic Repair of Test FlakinessYang Chen, Reyhaneh JabbarvandISSTA 2024 · 8 citations
- Shipwright: A Human-in-the-Loop System for Dockerfile RepairJordan Henkel, Denini Silva, Leopoldo Teixeira, Marcelo d'Amorim et al.ICSE 2021 · 31 citations
- NIODebugger: A Novel Approach to Repair Non-Idempotent-Outcome Tests with LLM-Based AgentKaiyao KeICSE 2025 · 3 citations
- Automatic Dockerfile Generation with Large Language ModelsJun Lyu, He Zhang, Yusong Yuan, Lanxin Yang et al.ICSE 2026
