Microsecond Consensus for Microsecond Applications
Marcos K. Aguilera, Naama Ben-David, Rachid Guerraoui, Virendra J. Marathe, Athanasios Xygkis, Igor Zablotchi
Abstract
We consider the problem of making apps fault-tolerant through replication, when apps operate at the microsecond scale, as in finance, embedded computing, and microservices apps. These apps need a replication scheme that also operates at the microsecond scale, otherwise replication becomes a burden. We propose Mu, a system that takes less than 1.3 microseconds to replicate a (small) request in memory, and less than a millisecond to fail-over the system - this cuts the replication and fail-over latencies of the prior systems by at least 61% and 90%. Mu implements bona fide state machine replication/consensus (SMR) with strong consistency for a generic app, but it really shines on microsecond apps, where even the smallest overhead is significant. To provide this performance, Mu introduces a new SMR protocol that carefully leverages RDMA. Roughly, in Mu a leader replicates a request by simply writing it directly to the log of other replicas using RDMA, without any additional communication. Doing so, however, introduces the challenge of handling concurrent leaders, changing leaders, garbage collecting the logs, and more - challenges that we address in this paper through a judicious combination of RDMA permissions and distributed algorithmic design. We implemented Mu and used it to replicate several systems: a financial exchange app called Liquibook, Redis, Memcached, and HERD. Our evaluation shows that Mu incurs a small replication latency, in some cases being the only viable replication system that incurs an acceptable overhead.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers24
- Rabia: Simplifying State-Machine Replication Through RandomizationHaochen Pan, Jesse Tuglu, Neo Zhou, Tianshu Wang et al.SOSP 2021 · 20 citations
- uBFT: Microsecond-Scale BFT using Disaggregated MemoryMarcos K. Aguilera, Naama Ben-David, Rachid Guerraoui, Antoine Murat et al.ASPLOS 2023 · 20 citations
- Resilient Baseband Processing in Virtualized RANs with SlingshotNikita Lazarev, Tao Ji, Anuj Kalia, Daehyeok Kim et al.SIGCOMM 2023 · 17 citations
- Halfmoon: Log-Optimal Fault-Tolerant Stateful Serverless ComputingSheng Qi, Xuanzhe Liu, Xin JinSOSP 2023 · 16 citations
- Waverunner: An Elegant Approach to Hardware Acceleration of State Machine ReplicationMohammadreza Alimadadi, Hieu Mai, Shenghsun Cho, Michael Ferdman et al.NSDI 2023 · 16 citations
Builds on2
- HovercRaft: achieving scalability and fault-tolerance for microsecond-scale datacenter servicesMarios Kogias, Edouard BugnionEuroSys 2020 · 52 citations
- Hermes: A Fast, Fault-Tolerant and Linearizable Replication ProtocolAntonios Katsarakis, Vasilis Gavrielatos, M. R. Siavash Katebzadeh, Arpit Joshi et al.ASPLOS 2020 · 47 citations
Related papers
- uKharon: A Membership Service for Microsecond ApplicationsRachid Guerraoui, Antoine Murat, Javier Picorel, Athanasios Xygkis et al.USENIX ATC 2022
- Bandle: Asynchronous State Machine Replication Made EfficientBo Wang, Shengyun Liu, He Dong, Xiangzhe Wang et al.EuroSys 2024 · 4 citations
- Odyssey: the impact of modern hardware on strongly-consistent replication protocolsVasilis Gavrielatos, Antonios Katsarakis, Vijay NagarajanEuroSys 2021 · 11 citations
- Rashnu: Data-Dependent Order-FairnessHeena Nagda, Shubhendra Pal Singhal, Mohammad Javad Amiri, Boon Thau LooVLDB 2024 · 9 citations
- Patronus: High-Performance and Protective Remote MemoryBin Yan, Youyou Lu, Qing Wang, Minhui Xie et al.FAST 2023
