State-Compute Replication: Parallelizing High-Speed Stateful Packet Processing
Qiongwen Xu, Sebastiano Miano, Xiangyu Gao, Tao Wang, Adithya Murugadass, Songyuan Zhang, Anirudh Sivaraman, Gianni Antichi, Srinivas Narayana
Abstract
With the slowdown of Moore's law, CPU-oriented packet processing in software will be significantly outpaced by emerging line speeds of network interface cards (NICs). Single-core packet-processing throughput has saturated.
We consider the problem of high-speed packet processing with multiple CPU cores. The key challenge is state-memory that multiple packets must read and update. The prevailing method to scale throughput with multiple cores involves state sharding, processing all packets that update the same state, e.g., flow, at the same core. However, given the skewed nature of realistic flow size distributions, this method is untenable, since total throughput is limited by single core performance.
This paper introduces state-compute replication, a principle to scale the throughput of a single stateful flow across multiple cores using replication. Our design leverages a packet history sequencer running on a NIC or top-of-the-rack switch to enable multiple cores to update state without explicit synchronization. Our experiments with realistic data center and wide-area Internet traces shows that state-compute replication can scale total packet-processing throughput linearly with cores, independent of flow size distributions, across a range of realistic packet-processing programs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3e8e6bc6-e97e-4364-bec1-44a534861f64Cited by top-tier papers4
- Enabling AI Network Cross-Layer Design and Operations with Arcadia: A Simulation Platform at ScaleZhaodong Wang, Satyajeet Singh Ahuja, Xu Zhang, Max Noormohammadpour et al.NSDI 2026 · 3 citations
- DDoS Detection at the Scale of One Hundred TbpsYunming Xiao, Xijun Luo, Youliang Jiang, Aike Wang et al.NSDI 2026 · 2 citations
- Mitigating CPU Frontend for Complex Data Plane ApplicationsYihan Dang, Hao Li, Ze Xia, Jiajun Luan et al.NSDI 2026
- SBB: Eliminating Centralized Bottlenecks in Userspace Network RuntimeKang Hu, Shuqi Dong, Chuandong Li, Ran Yi et al.OSDI 2026
Builds on8
- Sailfish: accelerating cloud-scale multi-tenant multi-service gateways with programmable switchesTian Pan, Nianbing Yu, Chenhao Jia, Jianwen Pi et al.SIGCOMM 2021 · 111 citations
- Enabling Programmable Transport Protocols in High-Speed NICsMina Tahmasbi Arashloo, Alexey Lavrov, Manya Ghobadi, Jennifer Rexford et al.NSDI 2020 · 96 citations
- The nanoPU: A Nanosecond Network Stack for DatacentersStephen Ibanez, Alex Mallery, Serhat Arslan, Theo Jepsen et al.OSDI 2021 · 74 citations
- FlexTOE: Flexible TCP Offload with Fine-Grained ParallelismRajath Shashidhara, Tim Stamler, Antoine Kaufmann, Simon PeterNSDI 2022 · 66 citations
- Debugging Transient Faults in Data Centers using Synchronized Network-wide Packet HistoriesPravein Govindan Kannan, Nishant Budhdev, Raj Joshi, Mun Choon ChanNSDI 2021 · 27 citations
Related papers
- ShRing: Networking with Shared Receive RingsBoris Pismenny, Adam Morrison, Dan TsafrirOSDI 2023 · 13 citations
- Disentangling the Dual Role of NIC Receive RingsBoris Pismenny, Adam Morrison, Dan TsafrirOSDI 2025 · 2 citations
- Twenty Years After: Hierarchical Core-Stateless Fair QueueingZhuolong Yu, Jingfeng Wu, Vladimir Braverman, Ion Stoica et al.NSDI 2021 · 45 citations
- Stateful multi-pipelined programmable switchesVishal ShrivastavSIGCOMM 2022 · 24 citations
- NFlow and MVT Abstractions for NFV ScalingZiyan Wu, Yang Zhang, Wendi Feng, Zhi-Li ZhangINFOCOM 2022 · 6 citations
