Scalable, NearZero Loss Disaster Recovery for Distributed Data Stores
Ahmed Alquraan, Alex Kogan, Virendra J. Marathe, Samer Al-Kiswany
摘要
This paper presents a new Disaster Recovery (DR) system, called Slogger, that differs from prior works in two principle ways: (i) Slogger enables DR for a linearizable distributed data store, and (ii) Slogger adopts the continuous backup approach that strives to maintain a tiny lag on the backup site relative to the primary site, thereby restricting the data loss window, due to disasters, to milliseconds. These goals pose a significant set of challenges related to consistency of the backup site's state, failures, and scalability. Slogger employs a combination of asynchronous log replication, intra-data center synchronized clocks, pipelining, batching, and a novel watermark service to address these challenges. Furthermore, Slogger is designed to be deployable as an "add-on" module in an existing distributed data store with few modifications to the original code base. Our evaluation, conducted on Slogger extensions to a 32-sharded version of LogCabin, an open source key-value store, shows that Slogger maintains a very small data loss window of 14.2 milliseconds which is near the optimal value in our evaluation setup. Moreover, Slogger reduces the length of the data loss window by 50% compared to incremental snapshotting technique without having any performance penalty on the primary data store. Furthermore, our experiments demonstrate that Slogger achieves our other goals of scalability, fault tolerance, and efficient failover to the backup data store when a disaster is declared at the primary data store.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- Rethinking Logging, Checkpoints, and Recovery for High-Performance Storage EnginesMichael Haubenschild, Caetano Sauer, Thomas Neumann, Viktor LeisSIGMOD 2020 · 被引用 43 次
- Asynchronous Prefix Recoverability for Fast Distributed StoresTianyu Li, Badrish Chandramouli, Jose M. Faleiro, Samuel Madden 等SIGMOD 2021 · 被引用 7 次
- Scalog: Seamless Reconfiguration and Total Order in a Scalable Shared LogCong Ding, David Chu, Evan Zhao, Xiang Li 等NSDI 2020 · 被引用 52 次
- LeaseGuard: Raft Leases Done RightA. Jesse Jiryu Davis, Murat Demirbas, Lingzhi DengSIGMOD 2026 · 被引用 2 次
- Sundial: Fault-tolerant Clock Synchronization for DatacentersYuliang Li, Gautam Kumar, Hema Hariharan, Hassan M. G. Wassel 等OSDI 2020 · 被引用 66 次
