HoliPaxos: Towards More Predictable Performance in State Machine Replication
Zhiying Liang, Vahab Jabrayilov, Abutalib Aghayev, Aleksey Charapko
Abstract
State machine replication (SMR) algorithms ensure redundancy in critical systems and, as a result, underpin fault-tolerant distributed databases. Good SMR protocol performance is essential for capacity planning and meeting desired performance objectives. However, many implementations of popular SMR algorithms, such as MultiPaxos and Raft, have issues that make their performance unpredictable. This unpredictability often arises from certain "bolt-on" additions to core protocols, such as external failure detectors and replication log compaction. In this paper, we argue that tighter integration of such traditionally ad-hoc mechanisms with the core replication protocols can stabilize performance, making the solutions more reliable and more accessible to accurate capacity planning. Moreover, we show that these integrations can be non-disruptive for the underlying consensus algorithm, resulting in systems that preserve the simplicity and safety of traditional single-leader consensus-based SMR. To that order, we integrate the failure and slowdown detectors inside the SMR and achieve better performance and faster fail-over under various network partitions and node slowdown events. We also illustrate that tight integration of replication log management, pruning, and snapshotting can reduce memory and CPU usage while avoiding performance fluctuations associated with traditional log compaction and cleanup approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 368c0405-d89f-415c-afde-1f8d43f48e15Builds on10
- A large scale analysis of hundreds of in-memory cache clusters at TwitterJuncheng Yang, Yao Yue, K. V. RashmiOSDI 2020 · 245 citations
- Understanding, Detecting and Localizing Partial Failures in Large System SoftwareChang Lou, Peng Huang, Scott SmithNSDI 2020 · 88 citations
- Metastable Failures in the WildLexiang Huang, Matthew Magnusson, Abishek Bangalore Muralikrishna, Salman Estyak et al.OSDI 2022 · 38 citations
- Perseus: A Fail-Slow Detection Framework for Cloud Storage SystemsRuiming Lu, Erci Xu, Yiming Zhang, Fengyi Zhu et al.FAST 2023 · 31 citations
- Toward a Generic Fault Tolerance Technique for Partial Network PartitioningMohammed Alfatafta, Basil Alkhatib, Ahmed Alquraan, Samer Al-KiswanyOSDI 2020 · 30 citations
Related papers
- Rabia: Simplifying State-Machine Replication Through RandomizationHaochen Pan, Jesse Tuglu, Neo Zhou, Tianshu Wang et al.SOSP 2021 · 20 citations
- Bandle: Asynchronous State Machine Replication Made EfficientBo Wang, Shengyun Liu, He Dong, Xiangzhe Wang et al.EuroSys 2024 · 4 citations
- LSM-Raft: Optimizing Raft for LSM-tree StoreXiaojian Zhang, Xinyu Tan, Shaoxu Song, Xiangdong Huang et al.SIGMOD 2026 · 1 citation
- Tolerating Slowdowns in Replicated State Machines using CopilotsKhiem Ngo, Siddhartha Sen, Wyatt LloydOSDI 2020 · 24 citations
- Scaling Replicated State Machines with CompartmentalizationMichael J. Whittaker, Ailidani Ailijiang, Aleksey Charapko, Murat Demirbas et al.VLDB 2021 · 40 citations
