Tolerating Slowdowns in Replicated State Machines using Copilots
Khiem Ngo, Siddhartha Sen, Wyatt Lloyd
Abstract
Replicated state machines are linearizable, fault-tolerant groups of replicas that are coordinated using a consensus algorithm. Copilot replication is the first 1-slowdown-tolerant consensus protocol: it delivers normal latency despite the slowdown of any 1 replica. Copilot uses two distinguished replicas-the pilot and copilot-to proactively add redundancy to all stages of processing a client's command. Copilot uses dependencies and deduplication to resolve potentially differing orderings proposed by the pilots. To avoid dependencies leading to either pilot being able to slow down the group, Copilot uses fast takeovers that allow a fast pilot to complete the ongoing work of a slow pilot. Copilot includes two optimizations-ping-pong batching and null dependency elimination-that improve its performance when there are 0 and 1 slow pilots respectively. Our evaluation of Copilot shows its performance is lower but competitive with Multi-Paxos and EPaxos when no replicas are slow. When a replica is slow, Copilot is the only protocol that avoids high latencies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f84032d9-def3-4691-b658-6c06f4c82009Cited by top-tier papers10
- Metastable Failures in the WildLexiang Huang, Matthew Magnusson, Abishek Bangalore Muralikrishna, Salman Estyak et al.OSDI 2022 · 38 citations
- PrestigeBFT: Revolutionizing View Changes in BFT Consensus Algorithms with Reputation MechanismsGengrui Zhang, Fei Pan, Sofia Tijanic, Hans-Arno JacobsenICDE 2024 · 21 citations
- One-Size-Fits-None: Understanding and Enhancing Slow-Fault Tolerance in Modern Distributed SystemsRuiming Lu, Yunchi Lu, Yuxuan Jiang, Guangtao Xue et al.NSDI 2025 · 14 citations
- QuePaxa: Escaping the tyranny of timeouts in consensusPasindu Tennage, Cristina Basescu, Lefteris Kokoris-Kogias, Ewa Syta et al.SOSP 2023 · 7 citations
- Regular Sequential Serializability and Regular Sequential ConsistencyJeffrey Helt, Matthew Burke, Amit Levy, Wyatt LloydSOSP 2021 · 4 citations
Builds on3
- Gryff: Unifying Consensus and Shared RegistersMatthew Burke, Audrey Cheng, Wyatt LloydNSDI 2020 · 29 citations
- Meaningful AvailabilityTamas Hauer, Philipp Hoffmann, John Lunney, Dan Ardelean et al.NSDI 2020 · 20 citations
- Millions of Tiny DatabasesMarc Brooker, Tao Chen, Fan PingNSDI 2020 · 9 citations
Related papers
- Avicenna: Masking Slowdowns in Replicated State Machines with Counterfactual EvaluationChristopher Hodsdon, Zijian Qin, Khiem Ngo, Siddhartha Sen et al.EuroSys 2026
- Ambulance: Saving BFT through RacingNeil Giridharan, Shubham Mishra, Lorenzo Alvisi, Natacha Crooks et al.OSDI 2026 · 1 citation
- SwiftPaxos: Fast Geo-Replicated State MachinesFedor Ryabinin, Alexey Gotsman, Pierre SutraNSDI 2024 · 19 citations
- Bandle: Asynchronous State Machine Replication Made EfficientBo Wang, Shengyun Liu, He Dong, Xiangzhe Wang et al.EuroSys 2024 · 4 citations
- HoliPaxos: Towards More Predictable Performance in State Machine ReplicationZhiying Liang, Vahab Jabrayilov, Abutalib Aghayev, Aleksey CharapkoVLDB 2025 · 2 citations
