SC2021Top-tier venue
Pinpointing crash-consistency bugs in the HPC I/O stack: a cross-layer approach
Jinghan Sun, Jian Huang, Marc Snir
2021Year
5Citations
Abstract
We present ParaCrash, a testing framework for studying crash recovery in a typical HPC I/O stack, and demonstrate its use by identifying 15 new crash-consistency bugs in various parallel file systems (PFS) and I/O libraries. ParaCrash uses a "golden version" approach to test the entire HPC I/O stack: storage state after recovery from a crash is correct if it matches the state that can be achieved by a partial execution with no crashes. It supports systematic testing of a multilayered I/O stack while properly identifying the layer responsible for the bugs.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- Chipmunk: Investigating Crash-Consistency in Persistent-Memory File SystemsHayley LeBlanc, Shankara Pailoor, Om Saran K. R. E., Isil Dillig et al.EuroSys 2023 · 11 citations
- File System Semantics Requirements of HPC ApplicationsChen Wang, Kathryn Mohror, Marc SnirHPDC 2021 · 25 citations
- Fast and Parallelized Crash Consistency with Opportunistic Order EliminationJiahao Chen, Yanqi Pan, Wen Xia, Hao Huang et al.EuroSys 2026
- Testing file system implementations on layered modelsDongjie Chen, Yanyan Jiang, Chang Xu, Xiaoxing Ma et al.ICSE 2020 · 6 citations
- Coverage Guided Fault Injection for Cloud SystemsYu Gao, Wensheng Dou, Dong Wang, Wenhan Feng et al.ICSE 2023 · 13 citations
