Understanding host network stack overheads
Qizhe Cai, Shubham Chaudhary, Midhul Vuppalapati, Jaehyun Hwang, Rachit Agarwal
Abstract
Traditional end-host network stacks are struggling to keep up with rapidly increasing datacenter access link bandwidths due to their unsustainable CPU overheads. Motivated by this, our community is exploring a multitude of solutions for future network stacks: from Linux kernel optimizations to partial hardware offload to clean-slate userspace stacks to specialized host network hardware. The design space explored by these solutions would benefit from a detailed understanding of CPU inefficiencies in existing network stacks.
This paper presents measurement and insights for Linux kernel network stack performance for 100Gbps access link bandwidths. Our study reveals that such high bandwidth links, coupled with relatively stagnant technology trends for other host resources (e.g., core speeds and count, cache sizes, NIC buffer sizes, etc.), mark a fundamental shift in host network stack bottlenecks. For instance, we find that a single core is no longer able to process packets at line rate, with data copy from kernel to application buffers at the receiver becoming the core performance bottleneck. In addition, increase in bandwidth-delay products have outpaced the increase in cache sizes, resulting in inefficient DMA pipeline between the NIC and the CPU. Finally, we find that traditional loosely-coupled design of network stack and CPU schedulers in existing operating systems becomes a limiting factor in scaling network stack performance across cores. Based on insights from our study, we discuss implications to design of future operating systems, network protocols, and host hardware.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 19397a2b-b410-4ffa-83d4-d302098312adCited by top-tier papers36
- Host Congestion ControlSaksham Agarwal, Arvind Krishnamurthy, Rachit AgarwalSIGCOMM 2023 · 47 citations
- Strata: Hierarchical Context Caching for Long Context Language Model ServingZhiqiang Xie, Ziyi Xu, Mark Zhao, Yuwei An et al.OSDI 2026 · 40 citations
- dcPIM: near-optimal proactive datacenter transportQizhe Cai, Mina Tahmasbi Arashloo, Rachit AgarwalSIGCOMM 2022 · 30 citations
- RingLeader: Efficiently Offloading Intra-Server Orchestration to NICsJiaxin Lin, Adney Cardoza, Tarannum Khan, Yeonju Ro et al.NSDI 2023 · 26 citations
- zIO: Accelerating IO-Intensive Applications with Transparent Zero-Copy IOTimothy Stamler, Deukyeon Hwang, Amanda Raybuck, Wei Zhang et al.OSDI 2022 · 21 citations
Builds on3
- Enabling Programmable Transport Protocols in High-Speed NICsMina Tahmasbi Arashloo, Alexey Lavrov, Manya Ghobadi, Jennifer Rexford et al.NSDI 2020 · 96 citations
- Reexamining Direct Cache Access to Optimize I/O Intensive Applications for Multi-hundred-gigabit NetworksAlireza Farshin, Amir Roozbeh, Gerald Q. Maguire Jr., Dejan KosticUSENIX ATC 2020 · 88 citations
- TCP ≈ RDMA: CPU-efficient Remote Storage Access with i10Jaehyun Hwang, Qizhe Cai, Ao Tang, Rachit AgarwalNSDI 2020 · 70 citations
Related papers
- Understanding Host Network Stack LatencyTianyu Zuo, Jaehyun Hwang, Ao Tang, Rachit Agarwal et al.SIGCOMM 2026
- High-throughput and Flexible Host Networking for Accelerated ComputingAthinagoras Skiadopoulos, Zhiqiang Xie, Mark Zhao, Qizhe Cai et al.OSDI 2024 · 11 citations
- Parallelizing packet processing in container overlay networksJiaxin Lei, Manish Munikar, Kun Suo, Hui Lu et al.EuroSys 2021 · 22 citations
- Opening Up Kernel-Bypass TCP StacksShinichi Awamoto, Michio HondaUSENIX ATC 2025 · 5 citations
- Towards μs tail latency and terabit ethernet: disaggregating the host network stackQizhe Cai, Midhul Vuppalapati, Jaehyun Hwang, Christos Kozyrakis et al.SIGCOMM 2022 · 20 citations
