ALTOCUMULUS: Scalable Scheduling for Nanosecond-Scale Remote Procedure Calls
Jiechen Zhao, Iris Uwizeyimana, Karthik Ganesan, Mark C. Jeffrey, Natalie Enright Jerger
摘要
Online services in modern datacenters use Remote Procedure Calls (RPCs) to communicate between different software layers. Despite RPCs using just a few small functions, inefficient RPC handling can cause delays to propagate across the system and degrade end-to-end performance. Prior work has reduced RPC processing time to less than 1 μs, which now shifts the bottleneck to the scheduling of RPCs. Existing RPC schedulers suffer from either high overheads, inability to effectively utilize high core-count CPUs or do not adaptively fit different traffic patterns. To address these shortcomings, we present ALTOCUMULUS, 1 a scalable, software-hardware codesign to schedule RPCs at nanosecond scales. ALTOCUMULUS provides a proactive scheduling scheme and low-overhead messaging mechanism on top of a decentralized user runtime. ALTOCUMULUS also offers direct access from the user space to a set of simple hardware primitives to quickly migrate long-latency RPCs. We evaluate ALTOCUMULUS with synthetic workloads and an end-to-end in-memory key-value store application under real-world traffic patterns. ALTOCUMULUS improves throughput by 1.3-24.6× under a 99 th percentile latency <300 μs and reduces tail latency by up to 15.8× on 16-core systems over current state-of-the-art software and hardware schedulers. For 256-core systems, integrating ALTOCUMULUS with either a hardware-optimized NIC or commodity PCIe NIC can improve throughput by 2.8× or 2.7×, respectively, under 99 th percentile latency <8.5 μs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- μManycore: A Cloud-Native CPU for Tail at ScaleJovan Stojkovic, Chunao Liu, Muhammad Shahbaz, Josep TorrellasISCA 2023 · 被引用 16 次
- LibPreemptible: Enabling Fast, Adaptive, and Hardware-Assisted User-Space SchedulingYueying Li, Nikita Lazarev, David Koufaty, Tenny Yin 等HPCA 2024 · 被引用 14 次
- NetClone: Fast, Scalable, and Dynamic Request Cloning for Microsecond-Scale RPCsGyuyeong KimSIGCOMM 2023 · 被引用 4 次
- HardHarvest: Hardware-Supported Core Harvesting for MicroservicesJovan Stojkovic, Chunao Liu, Muhammad Shahbaz, Josep TorrellasISCA 2025 · 被引用 4 次
- AccelFlow: Orchestrating an On-Package Ensemble of Fine-Grained Accelerators for MicroservicesJovan Stojkovic, Abraham Farrell, Zhangxiaowen Gong, Christopher J. Hughes 等HPCA 2026 · 被引用 2 次
它引用的顶会 Paper16
- Caladan: Mitigating Interference at Microsecond TimescalesJoshua Fried, Zhenyuan Ruan, Amy Ousterhout, Adam BelayOSDI 2020 · 被引用 213 次
- Accelerometer: Understanding Acceleration Opportunities for Data Center Overheads at HyperscaleAkshitha Sriraman, Abhishek DhanotiaASPLOS 2020 · 被引用 78 次
- The nanoPU: A Nanosecond Network Stack for DatacentersStephen Ibanez, Alex Mallery, Serhat Arslan, Theo Jepsen 等OSDI 2021 · 被引用 74 次
- Mitosis: Transparently Self-Replicating Page-Tables for Large-Memory MachinesReto Achermann, Ashish Panwar, Abhishek Bhattacharjee, Timothy Roscoe 等ASPLOS 2020 · 被引用 62 次
- ghOSt: Fast & Flexible User-Space Delegation of Linux SchedulingJack Tigar Humphries, Neel Natu, Ashwin Chaugule, Ofir Weisse 等SOSP 2021 · 被引用 60 次
相关 Paper
- Turbo: SmartNIC-enabled Dynamic Load Balancing of µs-scale RPCsHamed Seyedroudbari, Srikar Vanavasam, Alexandros DaglisHPCA 2023 · 被引用 12 次
- The NEBULA RPC-Optimized ArchitectureMark Sutherland, Siddharth Gupta, Babak Falsafi, Virendra J. Marathe 等ISCA 2020 · 被引用 29 次
- RpcNIC: Enabling Efficient Datacenter RPC Offloading on PCIe-attached SmartNICsJie Zhang, Hongjing Huang, Xuzheng Chen, Xiang Li 等HPCA 2025 · 被引用 6 次
- Ship Compute or Ship Data? Why Not Both?Jie You, Jingfeng Wu, Xin Jin, Mosharaf ChowdhuryNSDI 2021 · 被引用 25 次
- Cerebros: Evading the RPC Tax in DatacentersArash Pourhabibi Zarandi, Mark Sutherland, Alexandros Daglis, Babak FalsafiMICRO 2021 · 被引用 24 次
