Cooperative Concurrency Control for Write-Intensive Key-Value Workloads
Mark Sutherland, Babak Falsafi, Alexandros Daglis
Abstract
Key-Value Stores (KVS) are foundational infrastructure components for online services. Due to their latency-critical nature, today’s best-performing KVS contain a plethora of full-stack optimizations commonly targeting read-mostly, popularity-skewed workloads. Motivated by production studies showing the increased prevalence of write-intensive workloads, we break down the KVS workload space into four distinct classes, and argue that current designs are only sufficient for two of them. The reason is that KVS concurrency control protocols expose a fundamental tradeoff: avoiding synchronization by partitioning writes across threads is mandatory for high throughput, but necessarily creates load imbalance that grows with core count and write fraction. We break this tradeoff with C-4, a co-design between NIC hardware and KVS software that judiciously separates write requests into two classes: independent ones that can be balanced across threads, and dependent ones which must be queued. C-4 dynamically partitions independent writes with the NIC to increase the load balancing flexibility of current KVS designs, and adds a software layer to the KVS to compact dependent writes into batches. Our evaluation shows that for write-intensive workloads, C-4 reduces 99th% tail latency by 1.3−5× and improves throughput by up to 1.7×.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 031811ac-2038-4d14-be86-8fe6ff4b3a4dCited by top-tier papers4
- Turbo: SmartNIC-enabled Dynamic Load Balancing of µs-scale RPCsHamed Seyedroudbari, Srikar Vanavasam, Alexandros DaglisHPCA 2023 · 12 citations
- Fast, Flexible, and Practical Kernel ExtensionsKumar Kartikeya Dwivedi, Rishabh R. Iyer, Sanidhya KashyapSOSP 2024 · 7 citations
- Single-Address-Space FaaS with JordYuanlong Li, Atri Bhattacharyya, Madhur Kumar, Abhishek Bhattacharjee et al.ISCA 2025 · 2 citations
- Prefix Siphoning: Exploiting LSM-Tree Range Filters For Information DisclosureAdi Kaufman, Moshik Hershcovitch, Adam MorrisonUSENIX ATC 2023
Builds on11
- A large scale analysis of hundreds of in-memory cache clusters at TwitterJuncheng Yang, Yao Yue, K. V. RashmiOSDI 2020 · 245 citations
- The nanoPU: A Nanosecond Network Stack for DatacentersStephen Ibanez, Alex Mallery, Serhat Arslan, Theo Jepsen et al.OSDI 2021 · 74 citations
- Segcache: a memory-efficient and scalable in-memory key-value cache for small objectsJuncheng Yang, Yao Yue, Rashmi VinayakNSDI 2021 · 70 citations
- Overload Control for µs-scale RPCs with BreakwaterInho Cho, Ahmed Saeed, Joshua Fried, Seo Jin Park et al.OSDI 2020 · 61 citations
- RackSched: A Microsecond-Scale Scheduler for Rack-Scale ComputersHang Zhu, Kostis Kaffes, Zixu Chen, Zhenming Liu et al.OSDI 2020 · 58 citations
Related papers
- Enabling Low Tail Latency on Multicore Key-Value StoresLucas Lersch, Ivan Schreter, Ismail Oukid, Wolfgang LehnerVLDB 2020 · 30 citations
- Dynamic read & write optimization with TurtleKVTony Astolfi, Vidya Silai, Darby Huye, Lan Liu et al.VLDB 2026
- The benefits of general-purpose on-NIC memoryBoris Pismenny, Liran Liss, Adam Morrison, Dan TsafrirASPLOS 2022 · 28 citations
- Holistic and Automated Task Scheduling for Distributed LSM-tree-based StorageYuanming Ren, Siyuan Sheng, Zhang Cao, Yongkun Li et al.FAST 2026 · 1 citation
- P4KVS: A Role-Replica Separation Offloading Method to Achieve In-Network Consistency for KV Stores Based on P4 SwitchesHaojuan Li, Zongpu Zhang, Chenzhen Ye, Ruohan Tang et al.SIGMOD 2026
