Scalable, NearZero Loss Disaster Recovery for Distributed Data Stores
Ahmed Alquraan, Alex Kogan, Virendra J. Marathe, Samer Al-Kiswany
Abstract
This paper presents a new Disaster Recovery (DR) system, called Slogger, that differs from prior works in two principle ways: (i) Slogger enables DR for a linearizable distributed data store, and (ii) Slogger adopts the continuous backup approach that strives to maintain a tiny lag on the backup site relative to the primary site, thereby restricting the data loss window, due to disasters, to milliseconds. These goals pose a significant set of challenges related to consistency of the backup site's state, failures, and scalability. Slogger employs a combination of asynchronous log replication, intra-data center synchronized clocks, pipelining, batching, and a novel watermark service to address these challenges. Furthermore, Slogger is designed to be deployable as an "add-on" module in an existing distributed data store with few modifications to the original code base. Our evaluation, conducted on Slogger extensions to a 32-sharded version of LogCabin, an open source key-value store, shows that Slogger maintains a very small data loss window of 14.2 milliseconds which is near the optimal value in our evaluation setup. Moreover, Slogger reduces the length of the data loss window by 50% compared to incremental snapshotting technique without having any performance penalty on the primary data store. Furthermore, our experiments demonstrate that Slogger achieves our other goals of scalability, fault tolerance, and efficient failover to the backup data store when a disaster is declared at the primary data store.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 79d6a710-9f94-4336-b808-39e539a549ceCited by top-tier papers1
Ask how each one uses itRelated papers
- Rethinking Logging, Checkpoints, and Recovery for High-Performance Storage EnginesMichael Haubenschild, Caetano Sauer, Thomas Neumann, Viktor LeisSIGMOD 2020 · 43 citations
- Asynchronous Prefix Recoverability for Fast Distributed StoresTianyu Li, Badrish Chandramouli, Jose M. Faleiro, Samuel Madden et al.SIGMOD 2021 · 7 citations
- Scalog: Seamless Reconfiguration and Total Order in a Scalable Shared LogCong Ding, David Chu, Evan Zhao, Xiang Li et al.NSDI 2020 · 52 citations
- LeaseGuard: Raft Leases Done RightA. Jesse Jiryu Davis, Murat Demirbas, Lingzhi DengSIGMOD 2026 · 2 citations
- Sundial: Fault-tolerant Clock Synchronization for DatacentersYuliang Li, Gautam Kumar, Hema Hariharan, Hassan M. G. Wassel et al.OSDI 2020 · 66 citations
