Disentangling the Dual Role of NIC Receive Rings
Boris Pismenny, Adam Morrison, Dan Tsafrir
Abstract
CPUs parallelize packet processing across cores via per-core receive (Rx) rings, which are typically sized to absorb bursts with ≥1Ki entries by default. The combined I/O working set (packet buffers pointed to by all Rx rings) easily exceeds the LLC capacity, thus degrading performance due to high memory bandwidth pressure. Recent work has reduced the I/O working set size by sharing Rx rings among cores with the "shRing" system. But this approach suffers from a bottleneck under imbalanced loads, which are common.
We contend that the bottleneck stems from an unnecessary entanglement of two orthogonal producer-consumer structures: (1) memory allocation, where the core produces empty buffers that the NIC consumes to store packets; and (2) packet delivery, where the NIC produces incoming packets that the core consumes. We propose rxBisect, a new CPU-NIC interface that decouples these structures. RxBisect replaces each Rx ring with two separate rings corresponding to the two structures, allowing memory allocation to be performed independently of packet reception. RxBisect can thus pass empty buffers efficiently between cores upon imbalance, thereby eliminating the aforementioned bottleneck. We implement rxBisect with software emulation and find that it improves throughput by up to 20% and 37% relative to the state-of-theart (shRing) and state-of-the-practice (per-core Rx rings).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fdd3e9b1-2410-43d9-b677-c4d5af8f5853Cited by top-tier papers3
- Efficient and Flexible Datapaths for Fine-Grained Rack-Scale Interconnects with Elastic QPChenxingyu Zhao, Yibo Wu, Hongtao Zhang, Jaehong Min et al.SIGCOMM 2026 · 1 citation
- When DDIO Meets Page Coloring: Revisiting DDIO Performance with SepiaChangwoo Song, Sanghyun Kim, Jinhyeok Oh, Qizhe Cai et al.OSDI 2026
- Toasty: Speeding Up Network I/O with Cache-Warm BuffersPreeti, Nitish Bhat, Ashwin Kumar, Mythili VutukuruASPLOS 2026
Builds on18
- Caladan: Mitigating Interference at Microsecond TimescalesJoshua Fried, Zhenyuan Ruan, Amy Ousterhout, Adam BelayOSDI 2020 · 213 citations
- Understanding host network stack overheadsQizhe Cai, Shubham Chaudhary, Midhul Vuppalapati, Jaehyun Hwang et al.SIGCOMM 2021 · 150 citations
- Reexamining Direct Cache Access to Optimize I/O Intensive Applications for Multi-hundred-gigabit NetworksAlireza Farshin, Amir Roozbeh, Gerald Q. Maguire Jr., Dejan KosticUSENIX ATC 2020 · 88 citations
- Characterizing Off-path SmartNIC for Accelerating Distributed SystemsXingda Wei, Rongxin Cheng, Yuhan Yang, Rong Chen et al.OSDI 2023 · 68 citations
- Making Kernel Bypass Practical for the Cloud with JunctionJoshua Fried, Gohar Irfan Chaudhry, Enrique Saurez, Esha Choukse et al.NSDI 2024 · 57 citations
Related papers
- ShRing: Networking with Shared Receive RingsBoris Pismenny, Adam Morrison, Dan TsafrirOSDI 2023 · 13 citations
- TiNA: Tiered Network Buffer Architecture for Fast Networking in Chiplet-based CPUsSiddharth Agarwal, Tianchen Wang, Jinghan Huang, Saksham Agarwal et al.ASPLOS 2026 · 1 citation
- State-Compute Replication: Parallelizing High-Speed Stateful Packet ProcessingQiongwen Xu, Sebastiano Miano, Xiangyu Gao, Tao Wang et al.NSDI 2025
- DRack: A CXL-Disaggregated Rack Architecture to Boost Inter-Rack CommunicationXu Zhang, Ke Liu, Yuan Hui, Xiaolong Zheng et al.USENIX ATC 2025 · 5 citations
- BBQ: A Block-based Bounded Queue for Exchanging Data and ProfilingJiawei Wang, Diogo Behrens, Ming Fu, Lilith Oberhauser et al.USENIX ATC 2022
