Hermes: A Fast, Fault-Tolerant and Linearizable Replication Protocol
Antonios Katsarakis, Vasilis Gavrielatos, M. R. Siavash Katebzadeh, Arpit Joshi, Aleksandar Dragojevic, Boris Grot, Vijay Nagarajan
摘要
Today's datacenter applications are underpinned by datastores that are responsible for providing availability, consistency, and performance. For high availability in the presence of failures, these datastores replicate data across several nodes. This is accomplished with the help of a reliable replication protocol that is responsible for maintaining the replicas strongly-consistent even when faults occur. Strong consistency is preferred to weaker consistency models that cannot guarantee an intuitive behavior for the clients. Furthermore, to accommodate high demand at real-time latencies, datastores must deliver high throughput and low latency.
This work introduces Hermes 1 , a broadcast-based reliable replication protocol for in-memory datastores that provides both high throughput and low latency by enabling local reads and fully-concurrent fast writes at all replicas. Hermes couples logical timestamps with cache-coherence-inspired invalidations to guarantee linearizability, avoid write serialization at a centralized ordering point, resolve write conflicts locally at each replica (hence ensuring that writes never abort) and provide fault-tolerance via replayable writes. Our implementation of Hermes over an RDMA-enabled reliable datastore with five replicas shows that Hermes consistently achieves higher throughput than state-of-the-art RDMA-based reliable protocols (ZAB and CRAQ) across all write ratios while also significantly reducing tail latency. At 5% writes, the tail latency of Hermes is 3.6× lower than that of CRAQ and ZAB.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper34
- Microsecond Consensus for Microsecond ApplicationsMarcos K. Aguilera, Naama Ben-David, Rachid Guerraoui, Virendra J. Marathe 等OSDI 2020 · 被引用 73 次
- ReDMArk: Bypassing RDMA Security MechanismsBenjamin Rothenberger, Konstantin Taranov, Adrian Perrig, Torsten HoeflerUSENIX Security 2021 · 被引用 56 次
- Avocado: A Secure In-Memory Distributed Storage SystemMaurice Bailleu, Dimitra Giantsidi, Vasilis Gavrielatos, Do Le Quoc 等USENIX ATC 2021 · 被引用 39 次
- CliqueMap: productionizing an RMA-based distributed caching systemArjun Singhvi, Aditya Akella, Maggie Anderson, Rob Cauble 等SIGCOMM 2021 · 被引用 22 次
- Confidential Consortium Framework: Secure Multiparty Applications with Confidentiality, Integrity, and High AvailabilityHeidi Howard, Fritz Alder, Edward Ashton, Amaury Chamayou 等VLDB 2024 · 被引用 22 次
相关 Paper
- The LAW theorem: Local Reads and Linearizable Asynchronous ReplicationEmmanouil Giortamis, Antonios Katsarakis, Vasilis Gavrielatos, Pramod Bhatotia 等VLDB 2025 · 被引用 2 次
- Zeus: locality-aware distributed transactionsAntonios Katsarakis, Yijun Ma, Zhaowei Tan, Andrew Bainbridge 等EuroSys 2021 · 被引用 20 次
- Odyssey: the impact of modern hardware on strongly-consistent replication protocolsVasilis Gavrielatos, Antonios Katsarakis, Vijay NagarajanEuroSys 2021 · 被引用 11 次
- IONIA: High-Performance Replication for Modern Disk-based KV StoresYi Xu, Henry Zhu, Prashant Pandey, Alex Conway 等FAST 2024 · 被引用 13 次
- HovercRaft: achieving scalability and fault-tolerance for microsecond-scale datacenter servicesMarios Kogias, Edouard BugnionEuroSys 2020 · 被引用 52 次
