FiDe: Reliable and Fast Crash Failure Detection to Boost Datacenter Coordination
Davide Rovelli, Pavel Chuprikov, Philipp Berdesinski, Ali Pahlevan, Patrick Jahnke, Patrick Eugster
Abstract
Failure detection is one of the most fundamental primitives on which distributed fault tolerant services and applications rely to achieve liveness. Typical crash failure detectors resort to using timeouts that have to take into account the unpredictability in interaction times among remote processes, caused by resource contention in the network and in endhost processors. While modern (gray) failure detectors have improved in detecting a wide range of failures, the problem of prohibitively large and unreliable timeouts for crash failures still persists, hampering performance of both the failure detector themselves and modern µs-scale services sitting on top.
We propose a novel fully reliable failure-detector (FiDe) that can report the crash of a remote process in a datacenter within less than 30 µs (7.2× faster than the current state of the art) with extremely high reliability, thanks to a ground-up design which provides stable end-to-end process interactions. By reliably lowering worst-case crash detection time, FiDe enables a class of algorithms that can be used to boost coordination services even in the absence of failures. We devise two novel, FiDe-based, highly efficient consensus protocols and integrate them into a key-value store and a synchronization service, improving throughput by up to 2.23× and reducing latency down to 0.46×.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b992c6f1-71e5-41f7-9cca-fdefa4d90aedBuilds on12
- Caladan: Mitigating Interference at Microsecond TimescalesJoshua Fried, Zhenyuan Ruan, Amy Ousterhout, Adam BelayOSDI 2020 · 213 citations
- Understanding, Detecting and Localizing Partial Failures in Large System SoftwareChang Lou, Peng Huang, Scott SmithNSDI 2020 · 88 citations
- The Demikernel Datapath OS Architecture for Microsecond-scale Datacenter SystemsIrene Zhang, Amanda Raybuck, Pratyush Patel, Kirk Olynyk et al.SOSP 2021 · 83 citations
- The nanoPU: A Nanosecond Network Stack for DatacentersStephen Ibanez, Alex Mallery, Serhat Arslan, Theo Jepsen et al.OSDI 2021 · 74 citations
- Microsecond Consensus for Microsecond ApplicationsMarcos K. Aguilera, Naama Ben-David, Rachid Guerraoui, Virendra J. Marathe et al.OSDI 2020 · 73 citations
Related papers
- Live in the Express LanePatrick Jahnke, Vincent Riesop, Pierre-Louis Roman, Pavel Chuprikov et al.USENIX ATC 2021 · 7 citations
- Sundial: Fault-tolerant Clock Synchronization for DatacentersYuliang Li, Gautam Kumar, Hema Hariharan, Hassan M. G. Wassel et al.OSDI 2020 · 66 citations
- Fast, Transparent Filesystem Microkernel Recovery with AnankeJing Liu, Yifan Dai, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-DusseauFAST 2025 · 6 citations
- uKharon: A Membership Service for Microsecond ApplicationsRachid Guerraoui, Antoine Murat, Javier Picorel, Athanasios Xygkis et al.USENIX ATC 2022
- FC: Adaptive Atomic Commit via Failure DetectionHexiang Pan, Quang-Trung Ta, Meihui Zhang, Zhanhao Zhao et al.ICDE 2024
