Lynx: A SmartNIC-driven Accelerator-centric Architecture for Network Servers
Maroun Tork, Lina Maudlej, Mark Silberstein
摘要
This paper explores new opportunities afforded by the growing deployment of compute and I/O accelerators to improve the performance and efficiency of hardware-accelerated computing services in data centers.
We propose Lynx, an accelerator-centric network server architecture that offloads the server data and control planes to the SmartNIC, and enables direct networking from accelerators via a lightweight hardware-friendly I/O mechanism. Lynx enables the design of hardware-accelerated network servers that run without CPU involvement, freeing CPU cores and improving performance isolation for accelerated services. It is portable across accelerator architectures and allows the management of both local and remote accelerators, seamlessly scaling beyond a single physical machine.
We implement and evaluate Lynx on GPUs and the Intel Visual Compute Accelerator, as well as two SmartNIC architectures -one with an FPGA, and another with an 8core ARM processor. Compared to a traditional host-centric approach, Lynx achieves over 4× higher throughput for a GPU-centric face verification server, where it is used for GPU communications with an external database, and 25% higher throughput for a GPU-accelerated neural network inference service. For this workload, we show that a single SmartNIC may drive 4 local and 8 remote GPUs while achieving linear performance scaling without using the host CPU.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- Reexamining Direct Cache Access to Optimize I/O Intensive Applications for Multi-hundred-gigabit NetworksAlireza Farshin, Amir Roozbeh, Gerald Q. Maguire Jr., Dejan KosticUSENIX ATC 2020 · 被引用 88 次
- Serverless computing on heterogeneous computersDong Du, Qingyuan Liu, Xueqiang Jiang, Yubin Xia 等ASPLOS 2022 · 被引用 68 次
- FpgaNIC: An FPGA-based Versatile 100Gb SmartNIC for GPUsZeke Wang, Hongjing Huang, Jie Zhang, Fei Wu 等USENIX ATC 2022 · 被引用 58 次
- GPU-Initiated On-Demand High-Throughput Storage Access in the BaM System ArchitectureZaid Qureshi, Vikram Sharma Mailthody, Isaac Gelado, Seungwon Min 等ASPLOS 2023 · 被引用 48 次
- MemLiner: Lining up Tracing and Application for a Far-Memory-Friendly RuntimeChenxi Wang, Haoran Ma, Shi Liu, Yifan Qiao 等OSDI 2022 · 被引用 47 次
相关 Paper
- FlexDriver: a network driver for your acceleratorHaggai Eran, Maxim Fudim, Gabi Malka, Gal Shalom 等ASPLOS 2022 · 被引用 14 次
- HAL: Hardware-assisted Load Balancing for Energy-efficient SNIC-Host Cooperative ComputingJinghan Huang, Jiaqi Lou, Srikar Vanavasam, Xinhao Kong 等ISCA 2024 · 被引用 7 次
- A Generic and Efficient Communication Framework for Message-Level In-Network ComputingXinchen Wan, Luyang Li, Han Tian, Xudong Liao 等INFOCOM 2025 · 被引用 2 次
- RpcNIC: Enabling Efficient Datacenter RPC Offloading on PCIe-attached SmartNICsJie Zhang, Hongjing Huang, Xuzheng Chen, Xiang Li 等HPCA 2025 · 被引用 6 次
- TrainBox: An Extreme-Scale Neural Network Training Server Architecture by Systematically Balancing OperationsPyeongsu Park, Heetaek Jeong, Jangwoo KimMICRO 2020 · 被引用 11 次
