Beaver: Practical Partial Snapshots for Distributed Cloud Services
Liangcheng Yu, Xiao Zhang, Haoran Zhang, John Sonchack, Dan R. K. Ports, Vincent Liu
摘要
Distributed snapshots are a classic class of protocols used for capturing a causally consistent view of states across machines. Although effective, existing protocols presume an isolated universe of processes to snapshot and require instrumentation and coordination of all. This assumption does not match today's cloud services-it is not always practical to instrument all involved processes nor realistic to assume zero interaction of the machines of interest with the external world.
To bridge this gap, this paper presents Beaver, the first practical partial snapshot protocol that ensures causal consistency under external traffic interference. Beaver presents a unique design point that tightly couples its protocol with the regularities of the underlying data center environment. By exploiting the placement of software load balancers in public clouds and their associated communication pattern, Beaver not only requires minimal changes to today's data center operations but also eliminates any form of blocking to existing communication, thus incurring near-zero overhead to user traffic. We demonstrate the Beaver's effectiveness through extensive testbed experiments and novel use cases.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Nightcore: efficient and scalable serverless computing for latency-sensitive, interactive microservicesZhipeng Jia, Emmett WitchelASPLOS 2021 · 被引用 218 次
- Sailfish: accelerating cloud-scale multi-tenant multi-service gateways with programmable switchesTian Pan, Nianbing Yu, Chenhao Jia, Jianwen Pi 等SIGCOMM 2021 · 被引用 111 次
- A High-Speed Load-Balancer Design with Guaranteed Per-Connection-ConsistencyTom Barbette, Chen Tang, Haoran Yao, Dejan Kostic 等NSDI 2020 · 被引用 100 次
- Tiara: A Scalable and Efficient Hardware Acceleration Architecture for Stateful Layer-4 Load BalancingChaoliang Zeng, Layong Luo, Teng Zhang, Zilong Wang 等NSDI 2022 · 被引用 97 次
- Sundial: Fault-tolerant Clock Synchronization for DatacentersYuliang Li, Gautam Kumar, Hema Hariharan, Hassan M. G. Wassel 等OSDI 2020 · 被引用 66 次
相关 Paper
- EXIST: Enabling Extremely Efficient Intra-Service Tracing Observability in DatacentersXinkai Wang, Xiaofeng Hou, Chao Li, Yuancheng Li 等ASPLOS 2025 · 被引用 4 次
- HovercRaft: achieving scalability and fault-tolerance for microsecond-scale datacenter servicesMarios Kogias, Edouard BugnionEuroSys 2020 · 被引用 52 次
- Uno: A One-Stop Solution for Inter- and Intra-Data Center Congestion Control and Reliable ConnectivityTommaso Bonato, Sepehr Abdous, Abdul Kabbani, Ahmad Ghalayini 等SC 2025 · 被引用 7 次
- Scalable, NearZero Loss Disaster Recovery for Distributed Data StoresAhmed Alquraan, Alex Kogan, Virendra J. Marathe, Samer Al-KiswanyVLDB 2020 · 被引用 4 次
- HeatSnap: A Hot Page-Aware Continuous Snapshots System for Virtual Machines in Web InfrastructureKangyue Gao, Chuangyu Ouyang, Xinkui Zhao, Miao Ye 等WWW 2025 · 被引用 2 次
