NetClone: Fast, Scalable, and Dynamic Request Cloning for Microsecond-Scale RPCs
Gyuyeong Kim
Abstract
Spawning duplicate requests, called cloning, is a powerful technique to reduce tail latency by masking service-time variability. However, traditional client-based cloning is static and harmful to performance under high load, while a recent coordinator-based approach is slow and not scalable. Both approaches are insufficient to serve modern microsecond-scale Remote Procedure Calls (RPCs). To this end, we present NetClone, a request cloning system that performs cloning decisions dynamically within nanoseconds at scale. Rather than the client or the coordinator, NetClone performs request cloning in the network switch by leveraging the capability of programmable switch ASICs. Specifically, NetClone replicates requests based on server states and blocks redundant responses using request fingerprints in the switch data plane. To realize the idea while satisfying the strict hardware constraints, we address several technical challenges when designing a custom switch data plane. NetClone can be integrated with emerging in-network request schedulers like RackSched. We implement a NetClone prototype with an Intel Tofino switch and a cluster of commodity servers. Our experimental results show that NetClone can improve the tail latency of microsecond-scale RPCs for synthetic and real-world application workloads and is robust to various system conditions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b3807f0e-7572-4c52-a12d-1769fe8436f4Cited by top-tier papers3
- Pushing the Limits of In-Network Caching for Key-Value StoresGyuyeong KimNSDI 2025 · 9 citations
- Towards Optimal Rack-scale μs-level CPU Scheduling through In-Network Workload ShapingXudong Liao, Han Tian, Xinchen Wan, Chaoliang Zeng et al.USENIX ATC 2025 · 1 citation
- Switch: Asynchronous Metadata Updating for Distributed Storage with in-Network Data VisibilityJunru Li, Qing Wang, Zhe Yang, Shuo Liu et al.ICDE 2026
Builds on13
- TEA: Enabling State-Intensive Network Functions on Programmable SwitchesDaehyeok Kim, Zaoxing Liu, Yibo Zhu, Changhoon Kim et al.SIGCOMM 2020 · 121 citations
- Pegasus: Tolerating Skewed Workloads in Distributed Storage with In-Network Coherence DirectoriesJialin Li, Jacob Nelson, Ellis Michael, Xin Jin et al.OSDI 2020 · 96 citations
- The nanoPU: A Nanosecond Network Stack for DatacentersStephen Ibanez, Alex Mallery, Serhat Arslan, Theo Jepsen et al.OSDI 2021 · 74 citations
- Overload Control for µs-scale RPCs with BreakwaterInho Cho, Ahmed Saeed, Joshua Fried, Seo Jin Park et al.OSDI 2020 · 61 citations
- NetLock: Fast, Centralized Lock Management Using Programmable SwitchesZhuolong Yu, Yiwen Zhang, Vladimir Braverman, Mosharaf Chowdhury et al.SIGCOMM 2020 · 60 citations
Related papers
- RackSched: A Microsecond-Scale Scheduler for Rack-Scale ComputersHang Zhu, Kostis Kaffes, Zixu Chen, Zhenming Liu et al.OSDI 2020 · 58 citations
- Draconis: Network-Accelerated Scheduling for Microsecond-Scale WorkloadsSreeharsha Udayashankar, Ashraf Abdel-Hadi, Ali José Mashtizadeh, Samer Al-KiswanyEuroSys 2024 · 5 citations
- NetRPC: Enabling In-Network Computation in Remote Procedure CallsBohan Zhao, Wenfei Wu, Wei XuNSDI 2023 · 23 citations
- In-Network Leaderless Replication for Distributed Data StoresGyuyeong Kim, Wonjun LeeVLDB 2022 · 11 citations
- RingLeader: Efficiently Offloading Intra-Server Orchestration to NICsJiaxin Lin, Adney Cardoza, Tarannum Khan, Yeonju Ro et al.NSDI 2023 · 26 citations
