Not A DPU in Name Only! Unleashing RDMA-capable DPUs in Multi-Tenant Serverless Clouds with NADINO
Shixiong Qi, Songyu Zhang, K. K. Ramakrishnan, Diman Zad Tootaghaj, Hardik Soni, Puneet Sharma
摘要
Serverless computing promises enhanced resource efficiency and lower user costs, yet is burdened by a heavyweight, CPU-bound data plane. Prior efforts exploiting shared memory reduce overhead locally but fall short when scaling across nodes. Furthermore, serverless environments can have unpredictable and large-scale multi-tenancy, leading to contention for shared network resources. We present NADINO, a DPU-centric serverless data plane that reduces the CPU burden and enables efficient, zero-copy communication in multi-tenant serverless clouds. Despite the limited general-purpose processing capability of the DPU cores, NADINO strategically exploits the DPU's potential by (1) offloading data transmission to high-performance NIC cores via RDMA, combined with intra-node shared memory to eliminate data copies across nodes, and (2) enabling cross-processor (CPU-DPU) shared memory to eliminate redundant data movement, which overwhelms wimpy DPU cores. At the core of NADINO is the DPU-enabled network engine (DNE) - a lightweight reverse proxy that isolates RDMA resources from tenant functions, orchestrates inter-node RDMA flows, and enforces fairness under contention. To further reduce CPU involvement, NADINO performs early HTTP/TCP-to-RDMA transport conversion at the cloud ingress, bridging the protocol mismatch before client traffic enters the RDMA fabric, thus avoiding costly protocol translation along the critical path. We show that careful selection of RDMA primitives (i.e., two-sided instead of one-sided) significantly affects the zero-copy data plane. Our preliminary experimental results show that enabling DPU offloading in NADINO improves RPS by 20.9×. The latency is reduced by a factor of 21× in the best case, all the while saving up to 7 CPU cores, and only consuming two wimpy DPU cores.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- StreamBox: A Lightweight GPU SandBox for Serverless Inference WorkflowHao Wu, Yue Yu, Junxiao Deng, Shadi Ibrahim 等USENIX ATC 2024 · 被引用 21 次
- Efficient Data Passing for Serverless Inference Workflows: A GPU-Centric ApproachHao Wu, Yaochen Liu, Minchen Yu, Qizhen Weng 等EuroSys 2026
- No Provisioned Concurrency: Fast RDMA-codesigned Remote Fork for Serverless ComputingXingda Wei, Fangming Lu, Tianxia Wang, Jinyu Gu 等OSDI 2023 · 被引用 78 次
- Maximizing the Benefit of RDMA at End HostsXiaoliang Wang, Hexiang Song, Cam-Tu Nguyen, Dongxu Cheng 等INFOCOM 2021 · 被引用 8 次
- Serialization/Deserialization-free State Transfer in Serverless WorkflowsFangming Lu, Xingda Wei, Zhuobin Huang, Rong Chen 等EuroSys 2024 · 被引用 31 次
