Turbo: SmartNIC-enabled Dynamic Load Balancing of µs-scale RPCs
Hamed Seyedroudbari, Srikar Vanavasam, Alexandros Daglis
摘要
Online services are decomposed into fine-grained software components that communicate over the network using fine-grained Remote Procedure Calls (RPCs). Inter-server communication often exhibits patterns of wide RPC fan-outs between software tiers, raising the well-known tail at scale effect and necessitating mechanisms that curb long response tail latencies. When handling µs-scale RPCs, request distribution across the cores of multicore servers is a major determinant of the resulting tail latency. Software approaches for inter-core RPC balancing introduce considerable overheads, throttling a server's peak throughput. On the other hand, existing NIC-based hardware mechanisms ameliorate software and inter-core synchronization overheads, but result in inter-core load imbalance that leaves significant performance improvement headroom.
We introduce Turbo, a hardware on-NIC load-balancing mechanism that achieves near-optimal inter-core load distribution for the most fine-grained, light-tailed RPCs with service times of only a couple of µs. We implement Turbo on a programmable NIC and evaluate it on a range of different service time distributions and with a high-performance Key-Value store. Compared to hardware NIC-based mechanisms that statically spread load across cores, Turbo boosts throughput under a 99% latency Service Level Objective (SLO) of 30× the service time by up to 5×, and by up to 95× for a more aggressive 10× SLO target.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- A Cloud-Scale Characterization of Remote Procedure CallsKorakit Seemakhupt, Brent E. Stephens, Samira Manabi Khan, Sihang Liu 等SOSP 2023 · 被引用 31 次
- HAL: Hardware-assisted Load Balancing for Energy-efficient SNIC-Host Cooperative ComputingJinghan Huang, Jiaqi Lou, Srikar Vanavasam, Xinhao Kong 等ISCA 2024 · 被引用 7 次
- RpcNIC: Enabling Efficient Datacenter RPC Offloading on PCIe-attached SmartNICsJie Zhang, Hongjing Huang, Xuzheng Chen, Xiang Li 等HPCA 2025 · 被引用 6 次
- Dynamic Load Balancer in Intel Xeon Scalable Processor: Performance Analyses, Enhancements, and GuidelinesJiaqi Lou, Srikar Vanavasam, Yifan Yuan, Ren Wang 等ISCA 2025 · 被引用 4 次
- Re-architecting End-host Networking with CXL: Coherence, Memory, and OffloadingHouxiang Ji, Yifan Yuan, Yang Zhou, Ipoom Jeong 等MICRO 2025 · 被引用 3 次
它引用的顶会 Paper10
- Caladan: Mitigating Interference at Microsecond TimescalesJoshua Fried, Zhenyuan Ruan, Amy Ousterhout, Adam BelayOSDI 2020 · 被引用 213 次
- PANIC: A High-Performance Programmable NIC for Multi-tenant NetworksJiaxin Lin, Kiran Patel, Brent E. Stephens, Anirudh Sivaraman 等OSDI 2020 · 被引用 104 次
- The nanoPU: A Nanosecond Network Stack for DatacentersStephen Ibanez, Alex Mallery, Serhat Arslan, Theo Jepsen 等OSDI 2021 · 被引用 74 次
- Dagger: efficient and fast RPCs in cloud microservices with near-memory reconfigurable NICsNikita Lazarev, Shaojie Xiang, Neil Adit, Zhiru Zhang 等ASPLOS 2021 · 被引用 54 次
- When Idling is Ideal: Optimizing Tail-Latency for Heavy-Tailed Datacenter Workloads with PerséphoneHenri Maxime Demoulin, Joshua Fried, Isaac Pedisich, Marios Kogias 等SOSP 2021 · 被引用 39 次
相关 Paper
- The NEBULA RPC-Optimized ArchitectureMark Sutherland, Siddharth Gupta, Babak Falsafi, Virendra J. Marathe 等ISCA 2020 · 被引用 29 次
- ALTOCUMULUS: Scalable Scheduling for Nanosecond-Scale Remote Procedure CallsJiechen Zhao, Iris Uwizeyimana, Karthik Ganesan, Mark C. Jeffrey 等MICRO 2022 · 被引用 11 次
- Sassy: SmartNIC-Assisted Notification Delivery for μs-Scale RDMA WorkloadsHamed Seyedroudbari, Alexandros DaglisHPCA 2026
- Cerebros: Evading the RPC Tax in DatacentersArash Pourhabibi Zarandi, Mark Sutherland, Alexandros Daglis, Babak FalsafiMICRO 2021 · 被引用 24 次
- RingLeader: Efficiently Offloading Intra-Server Orchestration to NICsJiaxin Lin, Adney Cardoza, Tarannum Khan, Yeonju Ro 等NSDI 2023 · 被引用 26 次
