Pegasus: Tolerating Skewed Workloads in Distributed Storage with In-Network Coherence Directories
Jialin Li, Jacob Nelson, Ellis Michael, Xin Jin, Dan R. K. Ports
Abstract
High performance distributed storage systems face the challenge of load imbalance caused by skewed and dynamic workloads. This paper introduces Pegasus, a new storage system that leverages new-generation programmable switch ASICs to balance load across storage servers. Pegasus uses selective replication of the most popular objects in the data store to distribute load. Using a novel in-network coherence directory, the Pegasus switch tracks and manages the location of replicated objects. This allows it to achieve load-aware forwarding and dynamic rebalancing for replicated keys, while still guaranteeing data coherence and consistency. The Pegasus design is practical to implement as it stores only forwarding metadata in the switch data plane. The resulting system improves the throughput of a distributed in-memory key-value store by more than 10x under a latency SLO — results which hold across a large set of workloads with varying degrees of skew, read/write ratio, object sizes, and dynamism.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a346590c-468c-434f-af2c-469fb253fd27Cited by top-tier papers33
- dLoRA: Dynamically Orchestrating Requests and Adapters for LoRA LLM ServingBingyang Wu, Ruidong Zhu, Zili Zhang, Peng Sun et al.OSDI 2024 · 79 citations
- Concordia: Distributed Shared Memory with In-Network Cache CoherenceQing Wang, Youyou Lu, Erci Xu, Junru Li et al.FAST 2021 · 74 citations
- MIND: In-Network Memory Management for Disaggregated Data CentersSeung-Seob Lee, Yanpeng Yu, Yupeng Tang, Anurag Khandelwal et al.SOSP 2021 · 51 citations
- Bidl: A High-throughput, Low-latency Permissioned Blockchain Framework for Datacenter NetworksJi Qi, Xusheng Chen, Yunpeng Jiang, Jianyu Jiang et al.SOSP 2021 · 36 citations
- RedPlane: enabling fault-tolerant stateful in-switch applicationsDaehyeok Kim, Jacob Nelson, Dan R. K. Ports, Vyas Sekar et al.SIGCOMM 2021 · 33 citations
Builds on5
- A large scale analysis of hundreds of in-memory cache clusters at TwitterJuncheng Yang, Yao Yue, K. V. RashmiOSDI 2020 · 245 citations
- The CacheLib Caching Engine: Design and Experiences at ScaleBenjamin Berg, Daniel S. Berger, Sara McAllister, Isaac Grosof et al.OSDI 2020 · 145 citations
- HotRing: A Hotspot-Aware In-Memory Key-Value StoreJiqiang Chen, Liang Chen, Sheng Wang, Guoyun Zhu et al.FAST 2020 · 80 citations
- Harmonia: Near-Linear Scalability for Replicated Storage with In-Network Conflict DetectionHang Zhu, Zhihao Bai, Jialin Li, Ellis Michael et al.VLDB 2020 · 58 citations
- Prism: Proxies without the PainYutaro Hayakawa, Michio Honda, Douglas Santry, Lars EggertNSDI 2021 · 32 citations
Related papers
- Switch: Asynchronous Metadata Updating for Distributed Storage with in-Network Data VisibilityJunru Li, Qing Wang, Zhe Yang, Shuo Liu et al.ICDE 2026
- SwitchFS: Asynchronous Metadata Updates for Distributed Filesystems with In-Network CoordinationJingwei Xu, Mingkai Dong, Qiulin Tian, Ziyi Tian et al.EuroSys 2026 · 1 citation
- FarReach: Write-back Caching in Programmable SwitchesSiyuan Sheng, Huancheng Puyang, Qun Huang, Lu Tang et al.USENIX ATC 2023 · 22 citations
- FLAIR: Accelerating Reads with Consistency-Aware Network RoutingHatem Takruri, Ibrahim Kettaneh, Ahmed Alquraan, Samer Al-KiswanyNSDI 2020 · 20 citations
- P4DB - The Case for In-Network OLTPMatthias Jasny, Lasse Thostrup, Tobias Ziegler, Carsten BinnigSIGMOD 2022 · 18 citations
