HovercRaft: achieving scalability and fault-tolerance for microsecond-scale datacenter services
Marios Kogias, Edouard Bugnion
Abstract
Cloud platform services must simultaneously be scalable, meet low tail latency service-level objectives, and be resilient to a combination of software, hardware, and network failures. Replication plays a fundamental role in meeting both the scalability and the fault-tolerance requirement, but is subject to opposing requirements: (1) scalability is typically achieved by relaxing consistency; (2) fault-tolerance is typically achieved through the consistent replication of state machines. Adding nodes to a system can therefore either increase performance at the expense of consistency, or increase resiliency at the expense of performance.
We propose HovercRaft, a new approach by which adding nodes increases both the resilience and the performance of general-purpose state-machine replication. We achieve this through an extension of the Raft protocol that carefully eliminates CPU and I/O bottlenecks and load balances requests.
Our implementation uses state-of-the-art kernel-bypass techniques, datacenter transport protocols, and in-network programmability to deliver up to 1 million operations/second for clusters of up to 9 nodes, linear speedup over unreplicated configuration for selected workloads, and a 4× speedup for the YCSBE-E benchmark running on Redis over an unreplicated deployment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers21
- Electrode: Accelerating Distributed Protocols with eBPFYang Zhou, Zezhou Wang, Sowmya Dharanipragada, Minlan YuNSDI 2023 · 78 citations
- Microsecond Consensus for Microsecond ApplicationsMarcos K. Aguilera, Naama Ben-David, Rachid Guerraoui, Virendra J. Marathe et al.OSDI 2020 · 73 citations
- Virtual Consensus in DelosMahesh Balakrishnan, Jason Flinn, Chen Shen, Mihir Dharamshi et al.OSDI 2020 · 42 citations
- Efficient replication via timestamp stabilityVitor Enes, Carlos Baquero, Alexey Gotsman, Pierre SutraEuroSys 2021 · 21 citations
- Halfmoon: Log-Optimal Fault-Tolerant Stateful Serverless ComputingSheng Qi, Xuanzhe Liu, Xin JinSOSP 2023 · 16 citations
Related papers
- Abraxas: Throughput-Efficient Hybrid Asynchronous ConsensusErica Blum, Jonathan Katz, Julian Loss, Kartik Nayak et al.CCS 2023 · 11 citations
- Avicenna: Masking Slowdowns in Replicated State Machines with Counterfactual EvaluationChristopher Hodsdon, Zijian Qin, Khiem Ngo, Siddhartha Sen et al.EuroSys 2026
- Tolerating Slowdowns in Replicated State Machines using CopilotsKhiem Ngo, Siddhartha Sen, Wyatt LloydOSDI 2020 · 24 citations
- SwiftPaxos: Fast Geo-Replicated State MachinesFedor Ryabinin, Alexey Gotsman, Pierre SutraNSDI 2024 · 19 citations
- Integrating 2PC with Consensus for Fast ReplicationYan Chen, Xinyi Yu, Shengyun Liu, Ruofan Xiong et al.SIGCOMM 2026
