Learning From Software Failures: A Case Study at a National Space Research Center
Dharun Anandayuvaraj, Tanmay Singla, Zain Alabedin Haj Hammadeh, Andreas Lund, Alexandra Holloway, James C. Davis
Abstract
Software failures can have significant consequences, making learning from failures a critical aspect of software engineering. While software organizations are recommended to conduct postmortems to learn from failures, the effectiveness and adoption of these practices vary widely. Understanding how engineers gather, document, share, and apply lessons from failures is essential for improving software reliability and preventing recurring failures. High-reliability organizations (HROs) often develop software systems where failures carry catastrophic risks, requiring continuous learning practices to ensure reliability. These organizations provide a valuable setting to examine practices and challenges for learning from software failures. Such insight could help develop processes and tools to improve reliability and to prevent recurring failures. However, we lack in-depth industry perspectives on the practices and challenges of learning from failures.
To address this gap, we conducted a case study through 10 indepth interviews with research software engineers at a national space research center. We examine how they learn from failures: how they gather, document, share, and apply lessons learned. To assess the transferability of our findings, we include data from 5 additional interviews at other HROs. Our findings provide insight on how software engineers learn from failures in practice. To summarize our findings: (1) failure learning is informal, ad-hoc, and inconsistently integrated into SDLC; (2) recurring failures persist due to the absence of structured processes; and (3) key challenges, including time constraints, knowledge loss due to team turnover & fragmented documentation, and weak process enforcement, undermine efforts to systematically learn from failures. Our findings contribute to a deeper understanding of how software engineers learn from failures and offer guidance for improving failure management practices.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on5
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Recommending Root-Cause and Mitigation Steps for Cloud Incidents using Large Language ModelsToufique Ahmed, Supriyo Ghosh, Chetan Bansal, Thomas Zimmermann et al.ICSE 2023 · 93 citations
- An Interview Study on Third-Party Cyber Threat Hunting Processes in the U.S. Department of Homeland SecurityWilliam P. Maxam III, James C. DavisUSENIX Security 2024 · 14 citations
- FAIL: Analyzing Software Failures from the News Using LLMsDharun Anandayuvaraj, Matthew Campbell, Arav Tewari, James C. DavisASE 2024 · 3 citations
- On the Contents and Utility of IoT Cybersecurity GuidelinesJesse Chen, Dharun Anandayuvaraj, James C. Davis, Sazzadur RahamanFSE 2024 · 1 citation
Related papers
- Faulty or Ready? Handling Failures in Deep-Learning Computer Vision Models until Deployment: A Study of Practices, Challenges, and NeedsAgathe Balayn, Natasa Rikalo, Jie Yang, Alessandro BozzonCHI 2023 · 10 citations
- Trust in Collaborative Automation in High Stakes Software Engineering Work: A Case Study at NASADavid Gray Widder, Laura Dabbish, James D. Herbsleb, Alexandra Holloway et al.CHI 2021 · 18 citations
- A Theory of Scientific Programming EfficacyElizaveta Pertseva, Melinda Chang, Ulia Zaman, Michael CoblenzICSE 2024 · 4 citations
- Navigating the Testing of Evolving Deep Learning Systems: An Exploratory Interview StudyHanmo You, Zan Wang, Bin Lin, Junjie ChenICSE 2025 · 1 citation
- Software Architecture in Practice: Challenges and OpportunitiesZhiyuan Wan, Yun Zhang, Xin Xia, Yi Jiang et al.FSE 2023 · 29 citations
