ShRing: Networking with Shared Receive Rings
Boris Pismenny, Adam Morrison, Dan Tsafrir
Abstract
Multicore systems parallelize to accommodate incoming Ethernet traffic, allocating one receive (Rx) ring with ≥1Ki entries per core by default. This ring size is sufficient to absorb packet bursts of single-core workloads. But the combined size of all Rx buffers (pointed to by all Rx rings) can exceed the size of the last-level cache. We observe that, in this case, NIC and CPU memory accesses are increasingly served by main memory, which might incur nonnegligible overheads when scaling to hundreds of incoming gigabits per second.
To alleviate this problem, we propose "shRing," which shares each Rx ring among several cores when networking memory bandwidth consumption is high. ShRing thus adds software synchronization costs, but this overhead is offset by the smaller memory footprint. We show that, consequently, shRing increases the throughput of NFV workloads by up to 1.27x, and that it reduces their latency by up to 38x. The substantial latency reduction occurs when shRing shortens the per-packet processing time to a value smaller than the packet interarrival time, thereby preventing overload conditions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7b0f7185-161b-495f-8a2e-1be9d9cdffa4Cited by top-tier papers11
- Making Kernel Bypass Practical for the Cloud with JunctionJoshua Fried, Gohar Irfan Chaudhry, Enrique Saurez, Esha Choukse et al.NSDI 2024 · 57 citations
- CEIO: A Cache-Efficient Network I/O Architecture for NIC-CPU Data PathsBowen Liu, Xinyang Huang, Qijing Li, Zhuobin Huang et al.SIGCOMM 2025 · 10 citations
- A4: Microarchitecture-Aware LLC Management for Datacenter Servers with Emerging I/O DevicesHaneul Park, Jiaqi Lou, Sangjin Lee, Yifan Yuan et al.ISCA 2025 · 2 citations
- Disentangling the Dual Role of NIC Receive RingsBoris Pismenny, Adam Morrison, Dan TsafrirOSDI 2025 · 2 citations
- Achieving Wire-Latency Storage Systems by Exploiting Hardware ACKsQing Wang, Jiwu Shu, Jing Wang, Yuhao ZhangNSDI 2025 · 1 citation
Builds on13
- Caladan: Mitigating Interference at Microsecond TimescalesJoshua Fried, Zhenyuan Ruan, Amy Ousterhout, Adam BelayOSDI 2020 · 213 citations
- Understanding host network stack overheadsQizhe Cai, Shubham Chaudhary, Midhul Vuppalapati, Jaehyun Hwang et al.SIGCOMM 2021 · 150 citations
- Contention-Aware Performance Prediction For Virtualized Network FunctionsAntonis Manousis, Rahul Anand Sharma, Vyas Sekar, Justine SherrySIGCOMM 2020 · 57 citations
- Don't Forget the I/O When Allocating Your LLCYifan Yuan, Mohammad Alian, Yipeng Wang, Ren Wang et al.ISCA 2021 · 37 citations
- Autonomous NIC offloadsBoris Pismenny, Haggai Eran, Aviad Yehezkel, Liran Liss et al.ASPLOS 2021 · 32 citations
Related papers
- The benefits of general-purpose on-NIC memoryBoris Pismenny, Liran Liss, Adam Morrison, Dan TsafrirASPLOS 2022 · 28 citations
- State-Compute Replication: Parallelizing High-Speed Stateful Packet ProcessingQiongwen Xu, Sebastiano Miano, Xiangyu Gao, Tao Wang et al.NSDI 2025
- TiNA: Tiered Network Buffer Architecture for Fast Networking in Chiplet-based CPUsSiddharth Agarwal, Tianchen Wang, Jinghan Huang, Saksham Agarwal et al.ASPLOS 2026 · 1 citation
- RECANS: Low-Latency Network Function Chains with Hierarchical State SharingJian Zhao, Shujun Zhuang, Jian Li, Haibing GuanHPDC 2020 · 2 citations
- PeRF: Preemption-enabled RDMA FrameworkSugi Lee, Mingyu Choi, Ikjun Yeom, Younghoon KimUSENIX ATC 2024 · 3 citations
