Cooperative Concurrency Control for Write-Intensive Key-Value Workloads
Mark Sutherland, Babak Falsafi, Alexandros Daglis
摘要
Key-Value Stores (KVS) are foundational infrastructure components for online services. Due to their latency-critical nature, today’s best-performing KVS contain a plethora of full-stack optimizations commonly targeting read-mostly, popularity-skewed workloads. Motivated by production studies showing the increased prevalence of write-intensive workloads, we break down the KVS workload space into four distinct classes, and argue that current designs are only sufficient for two of them. The reason is that KVS concurrency control protocols expose a fundamental tradeoff: avoiding synchronization by partitioning writes across threads is mandatory for high throughput, but necessarily creates load imbalance that grows with core count and write fraction. We break this tradeoff with C-4, a co-design between NIC hardware and KVS software that judiciously separates write requests into two classes: independent ones that can be balanced across threads, and dependent ones which must be queued. C-4 dynamically partitions independent writes with the NIC to increase the load balancing flexibility of current KVS designs, and adds a software layer to the KVS to compact dependent writes into batches. Our evaluation shows that for write-intensive workloads, C-4 reduces 99th% tail latency by 1.3−5× and improves throughput by up to 1.7×.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Turbo: SmartNIC-enabled Dynamic Load Balancing of µs-scale RPCsHamed Seyedroudbari, Srikar Vanavasam, Alexandros DaglisHPCA 2023 · 被引用 12 次
- Fast, Flexible, and Practical Kernel ExtensionsKumar Kartikeya Dwivedi, Rishabh R. Iyer, Sanidhya KashyapSOSP 2024 · 被引用 7 次
- Single-Address-Space FaaS with JordYuanlong Li, Atri Bhattacharyya, Madhur Kumar, Abhishek Bhattacharjee 等ISCA 2025 · 被引用 2 次
- Prefix Siphoning: Exploiting LSM-Tree Range Filters For Information DisclosureAdi Kaufman, Moshik Hershcovitch, Adam MorrisonUSENIX ATC 2023
它引用的顶会 Paper11
- A large scale analysis of hundreds of in-memory cache clusters at TwitterJuncheng Yang, Yao Yue, K. V. RashmiOSDI 2020 · 被引用 245 次
- The nanoPU: A Nanosecond Network Stack for DatacentersStephen Ibanez, Alex Mallery, Serhat Arslan, Theo Jepsen 等OSDI 2021 · 被引用 74 次
- Segcache: a memory-efficient and scalable in-memory key-value cache for small objectsJuncheng Yang, Yao Yue, Rashmi VinayakNSDI 2021 · 被引用 70 次
- Overload Control for µs-scale RPCs with BreakwaterInho Cho, Ahmed Saeed, Joshua Fried, Seo Jin Park 等OSDI 2020 · 被引用 61 次
- RackSched: A Microsecond-Scale Scheduler for Rack-Scale ComputersHang Zhu, Kostis Kaffes, Zixu Chen, Zhenming Liu 等OSDI 2020 · 被引用 58 次
相关 Paper
- Enabling Low Tail Latency on Multicore Key-Value StoresLucas Lersch, Ivan Schreter, Ismail Oukid, Wolfgang LehnerVLDB 2020 · 被引用 30 次
- Dynamic read & write optimization with TurtleKVTony Astolfi, Vidya Silai, Darby Huye, Lan Liu 等VLDB 2026
- The benefits of general-purpose on-NIC memoryBoris Pismenny, Liran Liss, Adam Morrison, Dan TsafrirASPLOS 2022 · 被引用 28 次
- Holistic and Automated Task Scheduling for Distributed LSM-tree-based StorageYuanming Ren, Siyuan Sheng, Zhang Cao, Yongkun Li 等FAST 2026 · 被引用 1 次
- P4KVS: A Role-Replica Separation Offloading Method to Achieve In-Network Consistency for KV Stores Based on P4 SwitchesHaojuan Li, Zongpu Zhang, Chenzhen Ye, Ruohan Tang 等SIGMOD 2026
