XFir: Accelerating New-Flow Setup on Host Servers of a Large Cloud Network
Shihan Lin, Shunqiao Jiang, Liang Wang, Jian Wang, Chao Pei, Jian Zhao, Wenjun Wu, Kai Ren, Lijun Zhuang, Qingmin Liu, Heng Yu, Sirui Li
Abstract
In today's cloud networks, host servers widely deploy Data Processing Units (DPUs) as network accelerators under the "Sep-Path" paradigm. However, as server capabilities scale with increasing CPU cores and network bandwidth, the software slow path (executed on a DPU's CPU) has become a critical bottleneck for workloads with high new-flow rates. Meanwhile, new-flow setup logic on host servers must continuously evolve to meet diverse and changing customer demands, making flexibility a key requirement alongside performance. To address this gap, we present XFir, the first hardware-accelerated new-flow setup system for cloud host servers that delivers high CPS throughput while preserving sufficient flexibility. XFir leverages a next-generation DPU equipped with a Cloud Network co-Processor (CNP) to execute the host server's new-flow setup logic. XFir redesigns the host-server flow-setup datapath and table layout, optimizes LPM lookups, and introduces CPU-CNP collaboration mechanisms to further improve performance and reliability. Our evaluation shows that XFir achieves over 776K new-flow CPS on a single host server with 11.7μs slow-path latency. Compared to prior work (Fornax), XFir achieves 4.8x CPS and reduces latency by 69.2%. Moreover, XFir is cost-effective to deploy, requiring only a single DPU per host. Overall, XFir improves new-flow throughput while maintaining development flexibility at low financial cost.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers1
Ask how each one uses itRelated papers
- The nanoPU: A Nanosecond Network Stack for DatacentersStephen Ibanez, Alex Mallery, Serhat Arslan, Theo Jepsen et al.OSDI 2021 · 74 citations
- Analyzing Near-Network Hardware Acceleration with Co-Processing on DPUsDimitrios Giouroukis, Dwi P. A. Nugroho, Varun Pandey, Steffen Zeuch et al.VLDB 2025 · 3 citations
- CLIP: Accelerating Features Deployment for Programmable SwitchTingting Xu, Xiaoliang Wang, Chen Tian, Yun Xiong et al.INFOCOM 2023 · 6 citations
- dpKernels: Harvesting DPU Compute Resources for Data-path Efficiency in Cloud Data ProcessingJiasheng Hu, Kaiwen Zheng, Anna Li, Sidharth Sankhe et al.VLDB 2026
- Lynx: A SmartNIC-driven Accelerator-centric Architecture for Network ServersMaroun Tork, Lina Maudlej, Mark SilbersteinASPLOS 2020 · 64 citations
