Beaver: Practical Partial Snapshots for Distributed Cloud Services
Liangcheng Yu, Xiao Zhang, Haoran Zhang, John Sonchack, Dan R. K. Ports, Vincent Liu
Abstract
Distributed snapshots are a classic class of protocols used for capturing a causally consistent view of states across machines. Although effective, existing protocols presume an isolated universe of processes to snapshot and require instrumentation and coordination of all. This assumption does not match today's cloud services-it is not always practical to instrument all involved processes nor realistic to assume zero interaction of the machines of interest with the external world.
To bridge this gap, this paper presents Beaver, the first practical partial snapshot protocol that ensures causal consistency under external traffic interference. Beaver presents a unique design point that tightly couples its protocol with the regularities of the underlying data center environment. By exploiting the placement of software load balancers in public clouds and their associated communication pattern, Beaver not only requires minimal changes to today's data center operations but also eliminates any form of blocking to existing communication, thus incurring near-zero overhead to user traffic. We demonstrate the Beaver's effectiveness through extensive testbed experiments and novel use cases.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ee9444c9-7170-463c-a26a-111b76b4e8ffBuilds on15
- Nightcore: efficient and scalable serverless computing for latency-sensitive, interactive microservicesZhipeng Jia, Emmett WitchelASPLOS 2021 · 218 citations
- Sailfish: accelerating cloud-scale multi-tenant multi-service gateways with programmable switchesTian Pan, Nianbing Yu, Chenhao Jia, Jianwen Pi et al.SIGCOMM 2021 · 111 citations
- A High-Speed Load-Balancer Design with Guaranteed Per-Connection-ConsistencyTom Barbette, Chen Tang, Haoran Yao, Dejan Kostic et al.NSDI 2020 · 100 citations
- Tiara: A Scalable and Efficient Hardware Acceleration Architecture for Stateful Layer-4 Load BalancingChaoliang Zeng, Layong Luo, Teng Zhang, Zilong Wang et al.NSDI 2022 · 97 citations
- Sundial: Fault-tolerant Clock Synchronization for DatacentersYuliang Li, Gautam Kumar, Hema Hariharan, Hassan M. G. Wassel et al.OSDI 2020 · 66 citations
Related papers
- EXIST: Enabling Extremely Efficient Intra-Service Tracing Observability in DatacentersXinkai Wang, Xiaofeng Hou, Chao Li, Yuancheng Li et al.ASPLOS 2025 · 4 citations
- HovercRaft: achieving scalability and fault-tolerance for microsecond-scale datacenter servicesMarios Kogias, Edouard BugnionEuroSys 2020 · 52 citations
- Uno: A One-Stop Solution for Inter- and Intra-Data Center Congestion Control and Reliable ConnectivityTommaso Bonato, Sepehr Abdous, Abdul Kabbani, Ahmad Ghalayini et al.SC 2025 · 7 citations
- Scalable, NearZero Loss Disaster Recovery for Distributed Data StoresAhmed Alquraan, Alex Kogan, Virendra J. Marathe, Samer Al-KiswanyVLDB 2020 · 4 citations
- HeatSnap: A Hot Page-Aware Continuous Snapshots System for Virtual Machines in Web InfrastructureKangyue Gao, Chuangyu Ouyang, Xinkui Zhao, Miao Ye et al.WWW 2025 · 2 citations
