NetClone: Fast, Scalable, and Dynamic Request Cloning for Microsecond-Scale RPCs
Gyuyeong Kim
摘要
Spawning duplicate requests, called cloning, is a powerful technique to reduce tail latency by masking service-time variability. However, traditional client-based cloning is static and harmful to performance under high load, while a recent coordinator-based approach is slow and not scalable. Both approaches are insufficient to serve modern microsecond-scale Remote Procedure Calls (RPCs). To this end, we present NetClone, a request cloning system that performs cloning decisions dynamically within nanoseconds at scale. Rather than the client or the coordinator, NetClone performs request cloning in the network switch by leveraging the capability of programmable switch ASICs. Specifically, NetClone replicates requests based on server states and blocks redundant responses using request fingerprints in the switch data plane. To realize the idea while satisfying the strict hardware constraints, we address several technical challenges when designing a custom switch data plane. NetClone can be integrated with emerging in-network request schedulers like RackSched. We implement a NetClone prototype with an Intel Tofino switch and a cluster of commodity servers. Our experimental results show that NetClone can improve the tail latency of microsecond-scale RPCs for synthetic and real-world application workloads and is robust to various system conditions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Pushing the Limits of In-Network Caching for Key-Value StoresGyuyeong KimNSDI 2025 · 被引用 9 次
- Towards Optimal Rack-scale μs-level CPU Scheduling through In-Network Workload ShapingXudong Liao, Han Tian, Xinchen Wan, Chaoliang Zeng 等USENIX ATC 2025 · 被引用 1 次
- Switch: Asynchronous Metadata Updating for Distributed Storage with in-Network Data VisibilityJunru Li, Qing Wang, Zhe Yang, Shuo Liu 等ICDE 2026
它引用的顶会 Paper13
- TEA: Enabling State-Intensive Network Functions on Programmable SwitchesDaehyeok Kim, Zaoxing Liu, Yibo Zhu, Changhoon Kim 等SIGCOMM 2020 · 被引用 121 次
- Pegasus: Tolerating Skewed Workloads in Distributed Storage with In-Network Coherence DirectoriesJialin Li, Jacob Nelson, Ellis Michael, Xin Jin 等OSDI 2020 · 被引用 96 次
- The nanoPU: A Nanosecond Network Stack for DatacentersStephen Ibanez, Alex Mallery, Serhat Arslan, Theo Jepsen 等OSDI 2021 · 被引用 74 次
- Overload Control for µs-scale RPCs with BreakwaterInho Cho, Ahmed Saeed, Joshua Fried, Seo Jin Park 等OSDI 2020 · 被引用 61 次
- NetLock: Fast, Centralized Lock Management Using Programmable SwitchesZhuolong Yu, Yiwen Zhang, Vladimir Braverman, Mosharaf Chowdhury 等SIGCOMM 2020 · 被引用 60 次
相关 Paper
- RackSched: A Microsecond-Scale Scheduler for Rack-Scale ComputersHang Zhu, Kostis Kaffes, Zixu Chen, Zhenming Liu 等OSDI 2020 · 被引用 58 次
- Draconis: Network-Accelerated Scheduling for Microsecond-Scale WorkloadsSreeharsha Udayashankar, Ashraf Abdel-Hadi, Ali José Mashtizadeh, Samer Al-KiswanyEuroSys 2024 · 被引用 5 次
- NetRPC: Enabling In-Network Computation in Remote Procedure CallsBohan Zhao, Wenfei Wu, Wei XuNSDI 2023 · 被引用 23 次
- In-Network Leaderless Replication for Distributed Data StoresGyuyeong Kim, Wonjun LeeVLDB 2022 · 被引用 11 次
- RingLeader: Efficiently Offloading Intra-Server Orchestration to NICsJiaxin Lin, Adney Cardoza, Tarannum Khan, Yeonju Ro 等NSDI 2023 · 被引用 26 次
