FLEET: High-Performance Durable Replicated State Machines using Scattered and Coordinated Log Entries
Hua Fan, Hao Tan, Wenchao Zhou, Feifei Li
Abstract
Distributed coordination services are fundamental components of distributed systems, employing durable replicated state machines (RSMs) to ensure consistency across replicas and prevent data loss, even in the event of all nodes failing. These services typically rely on persistent logs for rapid recovery, as a universally agreed-upon log allows replicas to restore their state by sequentially replaying ordered log entries. However, the requirement for a totally ordered log inherently limits opportunities for parallelism.
This paper introduces Fleet, a high-performance durable RSM protocol that combines a hybrid scattered-entry log with an asynchronous ordered log. Our approach integrates synchronous persistence of scattered entries with asynchronous persistence of ordered entries, ensuring both rapid recovery and high levels of parallelism. Additionally, we propose a parallel applying optimization for the etcd database, named pre-apply. Experimental results demonstrate that Fleet significantly outperforms Raft and Scalog in terms of throughput and latency, achieving up to 10× the throughput under specific configurations and scaling effectively across multiple nodes. Additionally, with the pre-apply optimization, Fleet delivers a 10-fold increase in throughput compared to sequential applying on etcd. Although Fleet incurs a 5% overhead in recovery time during leader failure, this delay is tolerable given the rarity of such events.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d27e63f1-0d19-4a5d-89e0-ebb644ecdf7fBuilds on7
- Twine: A Unified Cluster Management System for Shared InfrastructureChunqiang Tang, Kenny Yu, Kaushik Veeraraghavan, Jonathan Kaldor et al.OSDI 2020 · 107 citations
- PigPaxos: Devouring the Communication Bottlenecks in Distributed ConsensusAleksey Charapko, Ailidani Ailijiang, Murat DemirbasSIGMOD 2021 · 55 citations
- Scalog: Seamless Reconfiguration and Total Order in a Scalable Shared LogCong Ding, David Chu, Evan Zhao, Xiang Li et al.NSDI 2020 · 52 citations
- Scaling Replicated State Machines with CompartmentalizationMichael J. Whittaker, Ailidani Ailijiang, Aleksey Charapko, Murat Demirbas et al.VLDB 2021 · 40 citations
- C5: Cloned Concurrency Control That Always Keeps UpJeffrey Helt, Abhinav Sharma, Daniel J. Abadi, Wyatt Lloyd et al.VLDB 2023 · 17 citations
Related papers
- LeaseGuard: Raft Leases Done RightA. Jesse Jiryu Davis, Murat Demirbas, Lingzhi DengSIGMOD 2026 · 2 citations
- HovercRaft: achieving scalability and fault-tolerance for microsecond-scale datacenter servicesMarios Kogias, Edouard BugnionEuroSys 2020 · 52 citations
- Asynchronous Prefix Recoverability for Fast Distributed StoresTianyu Li, Badrish Chandramouli, Jose M. Faleiro, Samuel Madden et al.SIGMOD 2021 · 7 citations
- MassBFT: Fast and Scalable Geo-Distributed Byzantine Fault-Tolerant ConsensusZeshun Peng, Yanfeng Zhang, Tinghao Feng, Weixing Zhou et al.ICDE 2025 · 1 citation
- The LAW theorem: Local Reads and Linearizable Asynchronous ReplicationEmmanouil Giortamis, Antonios Katsarakis, Vasilis Gavrielatos, Pramod Bhatotia et al.VLDB 2025 · 2 citations
