Sundial: Fault-tolerant Clock Synchronization for Datacenters
Yuliang Li, Gautam Kumar, Hema Hariharan, Hassan M. G. Wassel, Peter Hochschild, Dave Platt, Simon L. Sabato, Minlan Yu, Nandita Dukkipati, Prashant Chandra, Amin Vahdat
Abstract
Clock synchronization is critical for many datacenter applications such as distributed transactional databases, consistent snapshots, and network telemetry. As applications have increasing performance requirements and datacenter networks get into ultra-low latency, we need submicrosecond-level bound on time-uncertainty to reduce transaction delay and enable new network management applications (e.g., measuring one-way delay for congestion control). The state-of-the-art clock synchronization solutions focus on improving clock precision but may incur significant time-uncertainty bound due to the presence of failures. This significantly affects applications because in large-scale datacenters, temperature-related, link, device, and domain failures are common. We present Sundial, a fault-tolerant clock synchronization system for datacenters that achieves ∼100ns time-uncertainty bound under various types of failures. Sundial provides fast failure detection based on frequent synchronization messages in hardware. Sundial enables fast failure recovery using a novel graphbased algorithm to precompute a backup plan that is generic to failures. Through experiments in a >500-machine testbed and large-scale simulations, we show that Sundial can achieve ∼100ns time-uncertainty bound under different types of failures, which is more than two orders of magnitude lower than the state-of-the-art solutions. We also demonstrate the benefit of Sundial on applications such as Spanner and Swift congestion control.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d3d0006a-46db-4ae4-8c46-cfb489876e75Cited by top-tier papers24
- Aquila: A unified, low-latency fabric for datacenter networksDan Gibson, Hema Hariharan, Eric Lance, Moray McLaren et al.NSDI 2022 · 60 citations
- Poseidon: Efficient, Robust, and Practical Datacenter CC via Deployable INTWeitao Wang, Masoud Moshref, Yuliang Li, Gautam Kumar et al.NSDI 2023 · 58 citations
- Ekko: A Large-Scale Deep Learning Recommender System with Low-Latency Model UpdateChijun Sima, Yao Fu, Man-Kit Sit, Liyi Guo et al.OSDI 2022 · 31 citations
- Hydra: Serialization-Free Network Ordering for Strongly Consistent Distributed ApplicationsInho Choi, Ellis Michael, Yunfan Li, Dan R. K. Ports et al.NSDI 2023 · 21 citations
- Horus: Granular In-Network Task Scheduler for Cloud DatacentersParham Yassini, Khaled Diab, Saeed Mahloujifar, Mohamed HefeedaNSDI 2024 · 14 citations
Builds on1
Related papers
- Graham: Synchronizing Clocks by Leveraging Local Clock PropertiesAli Najafi, Michael WeiNSDI 2022
- Firefly: Scalable, Ultra-Accurate Clock Synchronization for DatacentersPooria Namyar, Yuliang Li, Weitao Wang, Nandita Dukkipati et al.SIGCOMM 2025 · 6 citations
- FiDe: Reliable and Fast Crash Failure Detection to Boost Datacenter CoordinationDavide Rovelli, Pavel Chuprikov, Philipp Berdesinski, Ali Pahlevan et al.USENIX ATC 2025 · 2 citations
- Live in the Express LanePatrick Jahnke, Vincent Riesop, Pierre-Louis Roman, Pavel Chuprikov et al.USENIX ATC 2021 · 7 citations
- MPulse: A Programmable and Autonomic Fault Detection System via Hierarchical Liveness ExchangeDi Wang, Haifeng Zhou, Zhengyan Zhou, Jiayu Luo et al.INFOCOM 2026
