Sundial: Fault-tolerant Clock Synchronization for Datacenters
Yuliang Li, Gautam Kumar, Hema Hariharan, Hassan M. G. Wassel, Peter Hochschild, Dave Platt, Simon L. Sabato, Minlan Yu, Nandita Dukkipati, Prashant Chandra, Amin Vahdat
摘要
Clock synchronization is critical for many datacenter applications such as distributed transactional databases, consistent snapshots, and network telemetry. As applications have increasing performance requirements and datacenter networks get into ultra-low latency, we need submicrosecond-level bound on time-uncertainty to reduce transaction delay and enable new network management applications (e.g., measuring one-way delay for congestion control). The state-of-the-art clock synchronization solutions focus on improving clock precision but may incur significant time-uncertainty bound due to the presence of failures. This significantly affects applications because in large-scale datacenters, temperature-related, link, device, and domain failures are common. We present Sundial, a fault-tolerant clock synchronization system for datacenters that achieves ∼100ns time-uncertainty bound under various types of failures. Sundial provides fast failure detection based on frequent synchronization messages in hardware. Sundial enables fast failure recovery using a novel graphbased algorithm to precompute a backup plan that is generic to failures. Through experiments in a >500-machine testbed and large-scale simulations, we show that Sundial can achieve ∼100ns time-uncertainty bound under different types of failures, which is more than two orders of magnitude lower than the state-of-the-art solutions. We also demonstrate the benefit of Sundial on applications such as Spanner and Swift congestion control.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper24
- Aquila: A unified, low-latency fabric for datacenter networksDan Gibson, Hema Hariharan, Eric Lance, Moray McLaren 等NSDI 2022 · 被引用 60 次
- Poseidon: Efficient, Robust, and Practical Datacenter CC via Deployable INTWeitao Wang, Masoud Moshref, Yuliang Li, Gautam Kumar 等NSDI 2023 · 被引用 58 次
- Ekko: A Large-Scale Deep Learning Recommender System with Low-Latency Model UpdateChijun Sima, Yao Fu, Man-Kit Sit, Liyi Guo 等OSDI 2022 · 被引用 31 次
- Hydra: Serialization-Free Network Ordering for Strongly Consistent Distributed ApplicationsInho Choi, Ellis Michael, Yunfan Li, Dan R. K. Ports 等NSDI 2023 · 被引用 21 次
- Horus: Granular In-Network Task Scheduler for Cloud DatacentersParham Yassini, Khaled Diab, Saeed Mahloujifar, Mohamed HefeedaNSDI 2024 · 被引用 14 次
它引用的顶会 Paper1
相关 Paper
- Graham: Synchronizing Clocks by Leveraging Local Clock PropertiesAli Najafi, Michael WeiNSDI 2022
- Firefly: Scalable, Ultra-Accurate Clock Synchronization for DatacentersPooria Namyar, Yuliang Li, Weitao Wang, Nandita Dukkipati 等SIGCOMM 2025 · 被引用 6 次
- FiDe: Reliable and Fast Crash Failure Detection to Boost Datacenter CoordinationDavide Rovelli, Pavel Chuprikov, Philipp Berdesinski, Ali Pahlevan 等USENIX ATC 2025 · 被引用 2 次
- Live in the Express LanePatrick Jahnke, Vincent Riesop, Pierre-Louis Roman, Pavel Chuprikov 等USENIX ATC 2021 · 被引用 7 次
- MPulse: A Programmable and Autonomic Fault Detection System via Hierarchical Liveness ExchangeDi Wang, Haifeng Zhou, Zhengyan Zhou, Jiayu Luo 等INFOCOM 2026
