HAL: Hardware-assisted Load Balancing for Energy-efficient SNIC-Host Cooperative Computing
Jinghan Huang, Jiaqi Lou, Srikar Vanavasam, Xinhao Kong, Houxiang Ji, Ipoom Jeong, Danyang Zhuo, Eun Kyung Lee, Nam Sung Kim
摘要
A typical SmartNIC (SNIC) integrates a processor comprising Arm CPU and accelerators with a conventional NIC. The processor is designed to energy-efficiently execute network functions frequently used by datacenter applications. With such a processor, the SNIC has promised to notably improve the system-wide energy efficiency of datacenter servers. Nevertheless, the latest trend of integrating accelerators into server CPUs for these functions sparks a question on the SNIC processor’s superiority over a host processor (i.e., server CPU with accelerators) in system-wide energy efficiency, especially under given tail latency constraints. Answering this question, we first take an Intel Xeon processor, integrated with various accelerators (e.g., QuickAssist Technology), as a host processor, and then compare it to an NVIDIA BlueField-2 SNIC processor. This uncovers that (1) the host accelerator, coupled with a more powerful memory subsystem, can outperform the SNIC accelerator, and (2) the SNIC processor can improve system-wide energy efficiency only at low packet rates for most functions under tail latency constraints. To provide high system-wide energy efficiency without compromising tail latency at any packet rates, we propose HAL, consisting of a hardware-based load balancer and an intelligent load balancing policy implemented inside the SNIC. When HAL determines that the SNIC processor cannot efficiently process a given function beyond a specific packet rate, it limits the rate of packets to the SNIC processor and lets the host processor handle the excess. We implement a HAL-enabled SNIC with a commodity FPGA and a BlueField-2 SNIC, plug it into a commodity server, and run 10 popular network functions. Our evaluation shows that HAL can improve the system-wide energy efficiency and throughput of the server running these functions by 31% and 10%, respectively, without notably increasing the tail latency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper12
- Serverless computing on heterogeneous computersDong Du, Qingyuan Liu, Xueqiang Jiang, Yubin Xia 等ASPLOS 2022 · 被引用 68 次
- Characterizing Off-path SmartNIC for Accelerating Distributed SystemsXingda Wei, Rongxin Cheng, Yuhan Yang, Rong Chen 等OSDI 2023 · 被引用 68 次
- Lynx: A SmartNIC-driven Accelerator-centric Architecture for Network ServersMaroun Tork, Lina Maudlej, Mark SilbersteinASPLOS 2020 · 被引用 64 次
- Xenic: SmartNIC-Accelerated Distributed TransactionsHenry N. Schuh, Weihao Liang, Ming Liu, Jacob Nelson 等SOSP 2021 · 被引用 62 次
- FpgaNIC: An FPGA-based Versatile 100Gb SmartNIC for GPUsZeke Wang, Hongjing Huang, Jie Zhang, Fei Wu 等USENIX ATC 2022 · 被引用 58 次
相关 Paper
- SmartNS: Enabling Line-rate and Flexible Network Stack with SmartNICXuzheng Chen, Jie Zhang, Baolin Zhu, Xueying Zhu 等EuroSys 2026
- Performance Prediction of On-NIC Network Functions with Multi-Resource Contention and Traffic AwarenessShaofeng Wu, Qiang Su, Zhixiong Niu, Hong XuASPLOS 2025 · 被引用 3 次
- Dynamic Load Balancer in Intel Xeon Scalable Processor: Performance Analyses, Enhancements, and GuidelinesJiaqi Lou, Srikar Vanavasam, Yifan Yuan, Ren Wang 等ISCA 2025 · 被引用 4 次
- RpcNIC: Enabling Efficient Datacenter RPC Offloading on PCIe-attached SmartNICsJie Zhang, Hongjing Huang, Xuzheng Chen, Xiang Li 等HPCA 2025 · 被引用 6 次
- Rambda: RDMA-driven Acceleration Framework for Memory-intensive µs-scale Datacenter ApplicationsYifan Yuan, Jinghan Huang, Yan Sun, Tianchen Wang 等HPCA 2023 · 被引用 25 次
