Speculative Recovery: Cheap, Highly Available Fault Tolerance with Disaggregated Storage
Nanqinqin Li, Anja Kalaba, Michael J. Freedman, Wyatt Lloyd, Amit Levy
Abstract
The ubiquity of disaggregated storage in cloud computing has led to a nascent technique for fault tolerance: instead of utilizing application-level replication, newly-launched backup instances recover application state from disaggregated storage (REDS) after a primary's failure. Attractively, REDS provides fault tolerance at a much lower cost than traditional replication schemes, wherein at least two instances are running. Failover in REDS is slow, however, because it sequentially first detects primary failure and only then starts recovery on a backup.
We propose speculative recovery to accelerate failover and thus increase the availability of applications using REDS. Instead of proceeding with failover sequentially, speculative recovery safely and efficiently parallelizes detecting primary failure and running recovery on a backup, by employing our new super and collapse primitives for disaggregated storage. Our implementation and evaluation of speculative recovery demonstrate that it considerably reduces failover time.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d4d85132-c41f-4d27-8aba-2d62267faf28Cited by top-tier papers3
- SplitFT: Fault Tolerance for Disaggregated Datacenters via Remote Memory LoggingXuhao Luo, Ramnatthan Alagappan, Aishwarya GanesanEuroSys 2024 · 2 citations
- Pilot Execution: Simulating Failure Recovery In Situ for Production Distributed SystemsZhenyu Li, Angting Cai, Chang LouNSDI 2026
- KUBEDIRECT: Unleashing the Full Power of the Cluster Manager for Serverless ComputingSheng Qi, Zhiquan Zhang, Xuanzhe Liu, Xin JinNSDI 2026
Builds on3
- LinnOS: Predictability on Unpredictable Flash Storage with a Light Neural NetworkMingzhe Hao, Levent Toksoz, Nanqinqin Li, Edward Edberg Halim et al.OSDI 2020 · 97 citations
- Gimbal: enabling multi-tenant storage disaggregation on SmartNIC JBOFsJaehong Min, Ming Liu, Tapan Chugh, Chenxingyu Zhao et al.SIGCOMM 2021 · 47 citations
- Rabia: Simplifying State-Machine Replication Through RandomizationHaochen Pan, Jesse Tuglu, Neo Zhou, Tianshu Wang et al.SOSP 2021 · 20 citations
Related papers
- SWARM: Replicating Shared Disaggregated-Memory Data in No TimeAntoine Murat, Clément Burgelin, Athanasios Xygkis, Igor Zablotchi et al.SOSP 2024 · 2 citations
- A Logically Disaggregated Cache for Replicated Storage SystemsKiran Hombal, Henry Zhu, Shreesha Gopalakrishna Bhat, Neil Kaushikkar et al.EuroSys 2026
- Scalable Distributed Inverted List Indexes in Disaggregated MemoryManuel Widmoser, Daniel Kocher, Nikolaus AugstenSIGMOD 2024 · 5 citations
- In-Storage Domain-Specific Acceleration for Serverless ComputingRohan Mahapatra, Soroush Ghodrati, Byung Hoon Ahn, Sean Kinzer et al.ASPLOS 2024 · 7 citations
- MTC: Scalable Transaction Commit for Multi-Primary Cloud DatabasesKecheng Luo, Xiaoxian Wei, Wenxin Liu, Peng Cai et al.ICDE 2026
