Rearchitecting Linux Storage Stack for µs Latency and High Throughput
Jaehyun Hwang, Midhul Vuppalapati, Simon Peter, Rachit Agarwal
Abstract
This paper demonstrates that it is possible to achieve µs-scale latency using Linux kernel storage stack, even when tens of latency-sensitive applications compete for host resources with throughput-bound applications that perform read/write operations at throughput close to hardware capacity. Furthermore, such performance can be achieved without any modification in applications, network hardware, kernel CPU schedulers and/or kernel network stack.
We demonstrate the above using design, implementation and evaluation of blk-switch, a new Linux kernel storage stack architecture. The key insight in blk-switch is that Linux's multi-queue storage design, along with multi-queue network and storage hardware, makes the storage stack conceptually similar to a network switch. blk-switch uses this insight to adapt techniques from the computer networking literature (e.g., multiple egress queues, prioritized processing of individual requests, load balancing, and switch scheduling) to the Linux kernel storage stack.
blk-switch evaluation over a variety of scenarios shows that it consistently achieves µs-scale average and tail latency (at both 99 th and 99.9 th percentiles), while allowing applications to near-perfectly utilize the hardware capacity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext febe0f79-8439-44e3-b74a-a4e467ba5fbbCited by top-tier papers25
- Empowering Azure Storage with RDMAWei Bai, Shanim Sainul Abdeen, Ankit Agrawal, Krishan Kumar Attre et al.NSDI 2023 · 117 citations
- XRP: In-Kernel Storage Functions with eBPFYuhong Zhong, Haoyu Li, Yu Jian Wu, Ioannis Zarkadas et al.OSDI 2022 · 100 citations
- Efficient Scheduling Policies for Microsecond-Scale TasksSarah McClure, Amy Ousterhout, Scott Shenker, Sylvia RatnasamyNSDI 2022 · 43 citations
- ODINFS: Scaling PM Performance with Opportunistic DelegationDiyu Zhou, Yuchen Qian, Vishal Gupta, Zhifei Yang et al.OSDI 2022 · 29 citations
- Karma: Resource Allocation for Dynamic DemandsMidhul Vuppalapati, Giannis Fikioris, Rachit Agarwal, Asaf Cidon et al.OSDI 2023 · 22 citations
Builds on3
- Caladan: Mitigating Interference at Microsecond TimescalesJoshua Fried, Zhenyuan Ruan, Amy Ousterhout, Adam BelayOSDI 2020 · 213 citations
- Building An Elastic Query Engine on Disaggregated StorageMidhul Vuppalapati, Justin Miron, Rachit Agarwal, Dan Truong et al.NSDI 2020 · 142 citations
- TCP ≈ RDMA: CPU-efficient Remote Storage Access with i10Jaehyun Hwang, Qizhe Cai, Ao Tang, Rachit AgarwalNSDI 2020 · 70 citations
Related papers
- SKQ: Event Scheduling for Optimizing Tail Latency in a Traditional OS KernelSiyao Zhao, Haoyu Gu, Ali José MashtizadehUSENIX ATC 2021 · 11 citations
- Towards μs tail latency and terabit ethernet: disaggregating the host network stackQizhe Cai, Midhul Vuppalapati, Jaehyun Hwang, Christos Kozyrakis et al.SIGCOMM 2022 · 20 citations
- Opening Up Kernel-Bypass TCP StacksShinichi Awamoto, Michio HondaUSENIX ATC 2025 · 5 citations
- Understanding Host Network Stack LatencyTianyu Zuo, Jaehyun Hwang, Ao Tang, Rachit Agarwal et al.SIGCOMM 2026
- λ-IO: A Unified IO Stack for Computational StorageZhe Yang, Youyou Lu, Xiaojian Liao, Youmin Chen et al.FAST 2023 · 54 citations
