Learning From Software Failures: A Case Study at a National Space Research Center
Dharun Anandayuvaraj, Tanmay Singla, Zain Alabedin Haj Hammadeh, Andreas Lund, Alexandra Holloway, James C. Davis
摘要
Software failures can have significant consequences, making learning from failures a critical aspect of software engineering. While software organizations are recommended to conduct postmortems to learn from failures, the effectiveness and adoption of these practices vary widely. Understanding how engineers gather, document, share, and apply lessons from failures is essential for improving software reliability and preventing recurring failures. High-reliability organizations (HROs) often develop software systems where failures carry catastrophic risks, requiring continuous learning practices to ensure reliability. These organizations provide a valuable setting to examine practices and challenges for learning from software failures. Such insight could help develop processes and tools to improve reliability and to prevent recurring failures. However, we lack in-depth industry perspectives on the practices and challenges of learning from failures.
To address this gap, we conducted a case study through 10 indepth interviews with research software engineers at a national space research center. We examine how they learn from failures: how they gather, document, share, and apply lessons learned. To assess the transferability of our findings, we include data from 5 additional interviews at other HROs. Our findings provide insight on how software engineers learn from failures in practice. To summarize our findings: (1) failure learning is informal, ad-hoc, and inconsistently integrated into SDLC; (2) recurring failures persist due to the absence of structured processes; and (3) key challenges, including time constraints, knowledge loss due to team turnover & fragmented documentation, and weak process enforcement, undermine efforts to systematically learn from failures. Our findings contribute to a deeper understanding of how software engineers learn from failures and offer guidance for improving failure management practices.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Recommending Root-Cause and Mitigation Steps for Cloud Incidents using Large Language ModelsToufique Ahmed, Supriyo Ghosh, Chetan Bansal, Thomas Zimmermann 等ICSE 2023 · 被引用 93 次
- An Interview Study on Third-Party Cyber Threat Hunting Processes in the U.S. Department of Homeland SecurityWilliam P. Maxam III, James C. DavisUSENIX Security 2024 · 被引用 14 次
- FAIL: Analyzing Software Failures from the News Using LLMsDharun Anandayuvaraj, Matthew Campbell, Arav Tewari, James C. DavisASE 2024 · 被引用 3 次
- On the Contents and Utility of IoT Cybersecurity GuidelinesJesse Chen, Dharun Anandayuvaraj, James C. Davis, Sazzadur RahamanFSE 2024 · 被引用 1 次
相关 Paper
- Faulty or Ready? Handling Failures in Deep-Learning Computer Vision Models until Deployment: A Study of Practices, Challenges, and NeedsAgathe Balayn, Natasa Rikalo, Jie Yang, Alessandro BozzonCHI 2023 · 被引用 10 次
- Trust in Collaborative Automation in High Stakes Software Engineering Work: A Case Study at NASADavid Gray Widder, Laura Dabbish, James D. Herbsleb, Alexandra Holloway 等CHI 2021 · 被引用 18 次
- A Theory of Scientific Programming EfficacyElizaveta Pertseva, Melinda Chang, Ulia Zaman, Michael CoblenzICSE 2024 · 被引用 4 次
- Navigating the Testing of Evolving Deep Learning Systems: An Exploratory Interview StudyHanmo You, Zan Wang, Bin Lin, Junjie ChenICSE 2025 · 被引用 1 次
- Software Architecture in Practice: Challenges and OpportunitiesZhiyuan Wan, Yun Zhang, Xin Xia, Yi Jiang 等FSE 2023 · 被引用 29 次
