FLEET: High-Performance Durable Replicated State Machines using Scattered and Coordinated Log Entries
Hua Fan, Hao Tan, Wenchao Zhou, Feifei Li
摘要
Distributed coordination services are fundamental components of distributed systems, employing durable replicated state machines (RSMs) to ensure consistency across replicas and prevent data loss, even in the event of all nodes failing. These services typically rely on persistent logs for rapid recovery, as a universally agreed-upon log allows replicas to restore their state by sequentially replaying ordered log entries. However, the requirement for a totally ordered log inherently limits opportunities for parallelism.
This paper introduces Fleet, a high-performance durable RSM protocol that combines a hybrid scattered-entry log with an asynchronous ordered log. Our approach integrates synchronous persistence of scattered entries with asynchronous persistence of ordered entries, ensuring both rapid recovery and high levels of parallelism. Additionally, we propose a parallel applying optimization for the etcd database, named pre-apply. Experimental results demonstrate that Fleet significantly outperforms Raft and Scalog in terms of throughput and latency, achieving up to 10× the throughput under specific configurations and scaling effectively across multiple nodes. Additionally, with the pre-apply optimization, Fleet delivers a 10-fold increase in throughput compared to sequential applying on etcd. Although Fleet incurs a 5% overhead in recovery time during leader failure, this delay is tolerable given the rarity of such events.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Twine: A Unified Cluster Management System for Shared InfrastructureChunqiang Tang, Kenny Yu, Kaushik Veeraraghavan, Jonathan Kaldor 等OSDI 2020 · 被引用 107 次
- PigPaxos: Devouring the Communication Bottlenecks in Distributed ConsensusAleksey Charapko, Ailidani Ailijiang, Murat DemirbasSIGMOD 2021 · 被引用 55 次
- Scalog: Seamless Reconfiguration and Total Order in a Scalable Shared LogCong Ding, David Chu, Evan Zhao, Xiang Li 等NSDI 2020 · 被引用 52 次
- Scaling Replicated State Machines with CompartmentalizationMichael J. Whittaker, Ailidani Ailijiang, Aleksey Charapko, Murat Demirbas 等VLDB 2021 · 被引用 40 次
- C5: Cloned Concurrency Control That Always Keeps UpJeffrey Helt, Abhinav Sharma, Daniel J. Abadi, Wyatt Lloyd 等VLDB 2023 · 被引用 17 次
相关 Paper
- LeaseGuard: Raft Leases Done RightA. Jesse Jiryu Davis, Murat Demirbas, Lingzhi DengSIGMOD 2026 · 被引用 2 次
- HovercRaft: achieving scalability and fault-tolerance for microsecond-scale datacenter servicesMarios Kogias, Edouard BugnionEuroSys 2020 · 被引用 52 次
- Asynchronous Prefix Recoverability for Fast Distributed StoresTianyu Li, Badrish Chandramouli, Jose M. Faleiro, Samuel Madden 等SIGMOD 2021 · 被引用 7 次
- MassBFT: Fast and Scalable Geo-Distributed Byzantine Fault-Tolerant ConsensusZeshun Peng, Yanfeng Zhang, Tinghao Feng, Weixing Zhou 等ICDE 2025 · 被引用 1 次
- The LAW theorem: Local Reads and Linearizable Asynchronous ReplicationEmmanouil Giortamis, Antonios Katsarakis, Vasilis Gavrielatos, Pramod Bhatotia 等VLDB 2025 · 被引用 2 次
