Understanding host network stack overheads
Qizhe Cai, Shubham Chaudhary, Midhul Vuppalapati, Jaehyun Hwang, Rachit Agarwal
摘要
Traditional end-host network stacks are struggling to keep up with rapidly increasing datacenter access link bandwidths due to their unsustainable CPU overheads. Motivated by this, our community is exploring a multitude of solutions for future network stacks: from Linux kernel optimizations to partial hardware offload to clean-slate userspace stacks to specialized host network hardware. The design space explored by these solutions would benefit from a detailed understanding of CPU inefficiencies in existing network stacks.
This paper presents measurement and insights for Linux kernel network stack performance for 100Gbps access link bandwidths. Our study reveals that such high bandwidth links, coupled with relatively stagnant technology trends for other host resources (e.g., core speeds and count, cache sizes, NIC buffer sizes, etc.), mark a fundamental shift in host network stack bottlenecks. For instance, we find that a single core is no longer able to process packets at line rate, with data copy from kernel to application buffers at the receiver becoming the core performance bottleneck. In addition, increase in bandwidth-delay products have outpaced the increase in cache sizes, resulting in inefficient DMA pipeline between the NIC and the CPU. Finally, we find that traditional loosely-coupled design of network stack and CPU schedulers in existing operating systems becomes a limiting factor in scaling network stack performance across cores. Based on insights from our study, we discuss implications to design of future operating systems, network protocols, and host hardware.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper36
- Host Congestion ControlSaksham Agarwal, Arvind Krishnamurthy, Rachit AgarwalSIGCOMM 2023 · 被引用 47 次
- Strata: Hierarchical Context Caching for Long Context Language Model ServingZhiqiang Xie, Ziyi Xu, Mark Zhao, Yuwei An 等OSDI 2026 · 被引用 40 次
- dcPIM: near-optimal proactive datacenter transportQizhe Cai, Mina Tahmasbi Arashloo, Rachit AgarwalSIGCOMM 2022 · 被引用 30 次
- RingLeader: Efficiently Offloading Intra-Server Orchestration to NICsJiaxin Lin, Adney Cardoza, Tarannum Khan, Yeonju Ro 等NSDI 2023 · 被引用 26 次
- zIO: Accelerating IO-Intensive Applications with Transparent Zero-Copy IOTimothy Stamler, Deukyeon Hwang, Amanda Raybuck, Wei Zhang 等OSDI 2022 · 被引用 21 次
它引用的顶会 Paper3
- Enabling Programmable Transport Protocols in High-Speed NICsMina Tahmasbi Arashloo, Alexey Lavrov, Manya Ghobadi, Jennifer Rexford 等NSDI 2020 · 被引用 96 次
- Reexamining Direct Cache Access to Optimize I/O Intensive Applications for Multi-hundred-gigabit NetworksAlireza Farshin, Amir Roozbeh, Gerald Q. Maguire Jr., Dejan KosticUSENIX ATC 2020 · 被引用 88 次
- TCP ≈ RDMA: CPU-efficient Remote Storage Access with i10Jaehyun Hwang, Qizhe Cai, Ao Tang, Rachit AgarwalNSDI 2020 · 被引用 70 次
相关 Paper
- Understanding Host Network Stack LatencyTianyu Zuo, Jaehyun Hwang, Ao Tang, Rachit Agarwal 等SIGCOMM 2026
- High-throughput and Flexible Host Networking for Accelerated ComputingAthinagoras Skiadopoulos, Zhiqiang Xie, Mark Zhao, Qizhe Cai 等OSDI 2024 · 被引用 11 次
- Parallelizing packet processing in container overlay networksJiaxin Lei, Manish Munikar, Kun Suo, Hui Lu 等EuroSys 2021 · 被引用 22 次
- Opening Up Kernel-Bypass TCP StacksShinichi Awamoto, Michio HondaUSENIX ATC 2025 · 被引用 5 次
- Towards μs tail latency and terabit ethernet: disaggregating the host network stackQizhe Cai, Midhul Vuppalapati, Jaehyun Hwang, Christos Kozyrakis 等SIGCOMM 2022 · 被引用 20 次
