Toward a Generic Fault Tolerance Technique for Partial Network Partitioning
Mohammed Alfatafta, Basil Alkhatib, Ahmed Alquraan, Samer Al-Kiswany
Abstract
We present an extensive study focused on partial network partitioning. Partial network partitions disrupt the communication between some but not all nodes in a cluster.
First, we conduct a comprehensive study of system failures caused by this fault in 12 popular systems. Our study reveals that the studied failures are catastrophic (e.g., lead to data loss), easily manifest, and can manifest by partially partitioning a single node.
Second, we dissect the design of eight popular systems and identify four principled approaches for tolerating partial partitions. Unfortunately, our analysis shows that implemented fault tolerance techniques are inadequate for modern systems; they either patch a particular mechanism or lead to a complete cluster shutdown, even when alternative network paths exist.
Finally, our findings motivate us to build Nifty, a transparent communication layer that masks partial network partitions. Nifty builds an overlay between nodes to detour packets around partial partitions. Our prototype evaluation with six popular systems shows that Nifty overcomes the shortcomings of current fault tolerance approaches and effectively masks partial partitions while imposing negligible overhead.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fbc104e1-a4f7-4a19-a37e-864e46123f31Cited by top-tier papers14
- Fail through the Cracks: Cross-System Interaction Failures in Modern Cloud SystemsLilia Tang, Chaitanya Bhandari, Yongle Zhang, Anna Karanika et al.EuroSys 2023 · 16 citations
- Efficient Exposure of Partial Failure Bugs in Distributed Systems with Inferred Abstract StatesHaoze Wu, Jia Pan, Peng HuangNSDI 2024 · 15 citations
- One-Size-Fits-None: Understanding and Enhancing Slow-Fault Tolerance in Modern Distributed SystemsRuiming Lu, Yunchi Lu, Yuxuan Jiang, Guangtao Xue et al.NSDI 2025 · 14 citations
- Understanding Transaction Bugs in Database SystemsZiyu Cui, Wensheng Dou, Yu Gao, Dong Wang et al.ICSE 2024 · 9 citations
- Demystifying and Checking Silent Semantic Violations in Large Distributed SystemsChang Lou, Yuzhuo Jing, Peng HuangOSDI 2022 · 8 citations
Related papers
- CoFI: Consistency-Guided Fault Injection for Cloud SystemsHaicheng Chen, Wensheng Dou, Dong Wang, Feng QinASE 2020 · 25 citations
- FAst in-network GraY failure detection for ISPsEdgar Costa Molero, Stefano Vissicchio, Laurent VanbeverSIGCOMM 2022 · 26 citations
- Jigsaw: A High-Utilization, Interference-Free Job Scheduler for Fat-Tree ClustersStaci A. Smith, David K. LowenthalHPDC 2021 · 4 citations
- MimicNet: fast performance estimates for data center networks with machine learningQizhen Zhang, Kelvin K. W. Ng, Charles W. Kazer, Shen Yan et al.SIGCOMM 2021 · 63 citations
- Understanding, Detecting and Localizing Partial Failures in Large System SoftwareChang Lou, Peng Huang, Scott SmithNSDI 2020 · 88 citations
