BMC: Accelerating Memcached using Safe In-kernel Caching and Pre-stack Processing
Yoann Ghigoff, Julien Sopena, Kahina Lazri, Antoine Blin, Gilles Muller
Abstract
In-memory key-value stores are critical components that help scale large internet services by providing low-latency access to popular data. Memcached, one of the most popular key-value stores, suffers from performance limitations inherent to the Linux networking stack and fails to achieve high performance when using high-speed network interfaces. While the Linux network stack can be bypassed using DPDK based solutions, such approaches require a complete redesign of the software stack and induce high CPU utilization even when client load is low.
To overcome these limitations, we present BMC, an inkernel cache for Memcached that serves requests before the execution of the standard network stack. Requests to the BMC cache are treated as part of the NIC interrupts, which allows performance to scale with the number of cores serving the NIC queues. To ensure safety, BMC is implemented using eBPF. Despite the safety constraints of eBPF, we show that it is possible to implement a complex cache service. Because BMC runs on commodity hardware and requires modification of neither the Linux kernel nor the Memcached application, it can be widely deployed on existing systems. BMC optimizes the processing time of Facebook-like small-size requests. On this target workload, our evaluations show that BMC improves throughput by up to 18x compared to the vanilla Memcached application and up to 6x compared to an optimized version of Memcached that uses the SO_REUSEPORT socket flag. In addition, our results also show that BMC has negligible overhead and does not deteriorate throughput when treating non-target workloads.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers27
- XRP: In-Kernel Storage Functions with eBPFYuhong Zhong, Haoyu Li, Yu Jian Wu, Ioannis Zarkadas et al.OSDI 2022 · 100 citations
- Electrode: Accelerating Distributed Protocols with eBPFYang Zhou, Zezhou Wang, Sowmya Dharanipragada, Minlan YuNSDI 2023 · 78 citations
- DINT: Fast In-Kernel Distributed Transactions with eBPFYang Zhou, Xingyu Xiang, Matthew Kiley, Sowmya Dharanipragada et al.NSDI 2024 · 40 citations
- Syrup: User-Defined Scheduling Across the StackKostis Kaffes, Jack Tigar Humphries, David Mazières, Christos KozyrakisSOSP 2021 · 35 citations
- LiteFlow: towards high-performance adaptive neural networks for kernel datapathJunxue Zhang, Chaoliang Zeng, Hong Zhang, Shuihai Hu et al.SIGCOMM 2022 · 23 citations
Builds on1
Related papers
- The benefits of general-purpose on-NIC memoryBoris Pismenny, Liran Liss, Adam Morrison, Dan TsafrirASPLOS 2022 · 28 citations
- Tux: Efficient Drop-in Networking for Database SystemsXinjing Zhou, Viktor Leis, Xiangyao Yu, Michael StonebrakerVLDB 2026 · 1 citation
- Put an Elephant into a Fridge: Optimizing Cache Efficiency for In-memory Key-value StoresKefei Wang, Jian Liu, Feng ChenVLDB 2020 · 23 citations
- DPA-Store: An Ordered Network Data Path Key-Value StoreFrederic Schimmelpfennig, Jan Sass, Reza Salkhordeh, Martin Kröning et al.OSDI 2026
- cache_ext: Customizing the Page Cache with eBPFTal Zussman, Ioannis Zarkadas, Jeremy Carin, Andrew Cheng et al.SOSP 2025 · 8 citations
