Fine-Grained Replicated State Machines for a Cluster Storage System
Ming Liu, Arvind Krishnamurthy, Harsha V. Madhyastha, Rishi Bhardwaj, Karan Gupta, Chinmay Kamat, Huapeng Yuan, Aditya Jaltade, Roger Liao, Pavan Konka, Anoop Jawahar
摘要
We describe the design and implementation of a consistent and fault-tolerant metadata index for a scalable block storage system. The block storage system supports the virtualized execution of legacy applications inside enterprise clusters by automatically distributing the stored blocks across the cluster's storage resources. To support the availability and scalability needs of the block storage system, we develop a distributed index that provides a replicated and consistent keyvalue storage abstraction.
The key idea underlying our design is the use of finegrained replicated state machines, wherein every key-value pair in the index is treated as a separate replicated state machine. This approach has many advantages over a traditional coarse-grained approach that represents an entire shard of data as a state machine: it enables effective use of multiple storage devices and cores, it is more robust to both short-and longterm skews in key access rates, and it can tolerate variations in key-value access latencies. The use of fine-grained replicated state machines, however, raises new challenges, which we address by co-designing the consensus protocol with the data store and streamlining the operation of the per-key replicated state machines. We demonstrate that fine-grained replicated state machines can provide significant performance benefits, characterize the performance of the system in the wild, and report on our experiences in building and deploying the system.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- LogNIC: A High-Level Performance Model for SmartNICsZerui Guo, Jiaxin Lin, Yuebin Bai, Daehyeok Kim 等MICRO 2023 · 被引用 17 次
- Odyssey: the impact of modern hardware on strongly-consistent replication protocolsVasilis Gavrielatos, Antonios Katsarakis, Vijay NagarajanEuroSys 2021 · 被引用 11 次
- Building an Elastic Block Storage over EBOFs Using Shadow ViewsSheng Jiang, Ming LiuNSDI 2025 · 被引用 10 次
- Understanding and Profiling NVMe-over-TCP Using ntprofYuyuan Kang, Ming LiuNSDI 2025 · 被引用 9 次
- Building Massive MIMO Baseband Processing on a Single-Node SupercomputerXincheng Xie, Wentao Hou, Zerui Guo, Ming LiuNSDI 2025 · 被引用 8 次
相关 Paper
- ResilientDB: Global Scale Resilient Blockchain FabricSuyash Gupta, Sajjad Rahnama, Jelle Hellings, Mohammad SadoghiVLDB 2020 · 被引用 100 次
- Avicenna: Masking Slowdowns in Replicated State Machines with Counterfactual EvaluationChristopher Hodsdon, Zijian Qin, Khiem Ngo, Siddhartha Sen 等EuroSys 2026
- Nezha: A Key-Value Separated Distributed Store with Optimized Raft IntegrationYangyang Wang, Yucong Dong, Ziqian Cheng, Zichen XuICDE 2026
- Hyra: Scalable Byzantine-Resilient State Storage Engine with Hierarchical Erasure-CodingQifeng Que, Xiaodong Qi, Zhao Zhang, Yanqin Yang 等SIGMOD 2026
- FaaSKeeper: Learning from Building Serverless Services with ZooKeeper as an ExampleMarcin Copik, Alexandru Calotoiu, Pengyu Zhou, Konstantin Taranov 等HPDC 2024 · 被引用 6 次
