Fine-Grained Replicated State Machines for a Cluster Storage System
Ming Liu, Arvind Krishnamurthy, Harsha V. Madhyastha, Rishi Bhardwaj, Karan Gupta, Chinmay Kamat, Huapeng Yuan, Aditya Jaltade, Roger Liao, Pavan Konka, Anoop Jawahar
Abstract
We describe the design and implementation of a consistent and fault-tolerant metadata index for a scalable block storage system. The block storage system supports the virtualized execution of legacy applications inside enterprise clusters by automatically distributing the stored blocks across the cluster's storage resources. To support the availability and scalability needs of the block storage system, we develop a distributed index that provides a replicated and consistent keyvalue storage abstraction.
The key idea underlying our design is the use of finegrained replicated state machines, wherein every key-value pair in the index is treated as a separate replicated state machine. This approach has many advantages over a traditional coarse-grained approach that represents an entire shard of data as a state machine: it enables effective use of multiple storage devices and cores, it is more robust to both short-and longterm skews in key access rates, and it can tolerate variations in key-value access latencies. The use of fine-grained replicated state machines, however, raises new challenges, which we address by co-designing the consensus protocol with the data store and streamlining the operation of the per-key replicated state machines. We demonstrate that fine-grained replicated state machines can provide significant performance benefits, characterize the performance of the system in the wild, and report on our experiences in building and deploying the system.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- LogNIC: A High-Level Performance Model for SmartNICsZerui Guo, Jiaxin Lin, Yuebin Bai, Daehyeok Kim et al.MICRO 2023 · 17 citations
- Odyssey: the impact of modern hardware on strongly-consistent replication protocolsVasilis Gavrielatos, Antonios Katsarakis, Vijay NagarajanEuroSys 2021 · 11 citations
- Building an Elastic Block Storage over EBOFs Using Shadow ViewsSheng Jiang, Ming LiuNSDI 2025 · 10 citations
- Understanding and Profiling NVMe-over-TCP Using ntprofYuyuan Kang, Ming LiuNSDI 2025 · 9 citations
- Building Massive MIMO Baseband Processing on a Single-Node SupercomputerXincheng Xie, Wentao Hou, Zerui Guo, Ming LiuNSDI 2025 · 8 citations
Related papers
- ResilientDB: Global Scale Resilient Blockchain FabricSuyash Gupta, Sajjad Rahnama, Jelle Hellings, Mohammad SadoghiVLDB 2020 · 100 citations
- Avicenna: Masking Slowdowns in Replicated State Machines with Counterfactual EvaluationChristopher Hodsdon, Zijian Qin, Khiem Ngo, Siddhartha Sen et al.EuroSys 2026
- Nezha: A Key-Value Separated Distributed Store with Optimized Raft IntegrationYangyang Wang, Yucong Dong, Ziqian Cheng, Zichen XuICDE 2026
- Hyra: Scalable Byzantine-Resilient State Storage Engine with Hierarchical Erasure-CodingQifeng Que, Xiaodong Qi, Zhao Zhang, Yanqin Yang et al.SIGMOD 2026
- FaaSKeeper: Learning from Building Serverless Services with ZooKeeper as an ExampleMarcin Copik, Alexandru Calotoiu, Pengyu Zhou, Konstantin Taranov et al.HPDC 2024 · 6 citations
