HoliPaxos: Towards More Predictable Performance in State Machine Replication
Zhiying Liang, Vahab Jabrayilov, Abutalib Aghayev, Aleksey Charapko
摘要
State machine replication (SMR) algorithms ensure redundancy in critical systems and, as a result, underpin fault-tolerant distributed databases. Good SMR protocol performance is essential for capacity planning and meeting desired performance objectives. However, many implementations of popular SMR algorithms, such as MultiPaxos and Raft, have issues that make their performance unpredictable. This unpredictability often arises from certain "bolt-on" additions to core protocols, such as external failure detectors and replication log compaction. In this paper, we argue that tighter integration of such traditionally ad-hoc mechanisms with the core replication protocols can stabilize performance, making the solutions more reliable and more accessible to accurate capacity planning. Moreover, we show that these integrations can be non-disruptive for the underlying consensus algorithm, resulting in systems that preserve the simplicity and safety of traditional single-leader consensus-based SMR. To that order, we integrate the failure and slowdown detectors inside the SMR and achieve better performance and faster fail-over under various network partitions and node slowdown events. We also illustrate that tight integration of replication log management, pruning, and snapshotting can reduce memory and CPU usage while avoiding performance fluctuations associated with traditional log compaction and cleanup approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- A large scale analysis of hundreds of in-memory cache clusters at TwitterJuncheng Yang, Yao Yue, K. V. RashmiOSDI 2020 · 被引用 245 次
- Understanding, Detecting and Localizing Partial Failures in Large System SoftwareChang Lou, Peng Huang, Scott SmithNSDI 2020 · 被引用 88 次
- Metastable Failures in the WildLexiang Huang, Matthew Magnusson, Abishek Bangalore Muralikrishna, Salman Estyak 等OSDI 2022 · 被引用 38 次
- Perseus: A Fail-Slow Detection Framework for Cloud Storage SystemsRuiming Lu, Erci Xu, Yiming Zhang, Fengyi Zhu 等FAST 2023 · 被引用 31 次
- Toward a Generic Fault Tolerance Technique for Partial Network PartitioningMohammed Alfatafta, Basil Alkhatib, Ahmed Alquraan, Samer Al-KiswanyOSDI 2020 · 被引用 30 次
相关 Paper
- Rabia: Simplifying State-Machine Replication Through RandomizationHaochen Pan, Jesse Tuglu, Neo Zhou, Tianshu Wang 等SOSP 2021 · 被引用 20 次
- Bandle: Asynchronous State Machine Replication Made EfficientBo Wang, Shengyun Liu, He Dong, Xiangzhe Wang 等EuroSys 2024 · 被引用 4 次
- LSM-Raft: Optimizing Raft for LSM-tree StoreXiaojian Zhang, Xinyu Tan, Shaoxu Song, Xiangdong Huang 等SIGMOD 2026 · 被引用 1 次
- Tolerating Slowdowns in Replicated State Machines using CopilotsKhiem Ngo, Siddhartha Sen, Wyatt LloydOSDI 2020 · 被引用 24 次
- Scaling Replicated State Machines with CompartmentalizationMichael J. Whittaker, Ailidani Ailijiang, Aleksey Charapko, Murat Demirbas 等VLDB 2021 · 被引用 40 次
