The nanoPU: A Nanosecond Network Stack for Datacenters
Stephen Ibanez, Alex Mallery, Serhat Arslan, Theo Jepsen, Muhammad Shahbaz, Changhoon Kim, Nick McKeown
Abstract
We present the nanoPU, a new NIC-CPU co-design to accelerate an increasingly pervasive class of datacenter applications: those that utilize many small Remote Procedure Calls (RPCs) with very short (µs-scale) processing times. The novel aspect of the nanoPU is the design of a fast path between the network and applications-bypassing the cache and memory hierarchy, and placing arriving messages directly into the CPU register file. This fast path contains programmable hardware support for low latency transport and congestion control as well as hardware support for efficient load balancing of RPCs to cores. A hardware-accelerated thread scheduler makes subnanosecond decisions, leading to high CPU utilization and low tail response time for RPCs.
We built an FPGA prototype of the nanoPU fast path by modifying an open-source RISC-V CPU, and evaluated its performance using cycle-accurate simulations on AWS FPGAs. The wire-to-wire RPC response time through the nanoPU is just 69ns, an order of magnitude quicker than the best-ofbreed, low latency, commercial NICs. We demonstrate that the hardware thread scheduler is able to lower RPC tail response time by about 5× while enabling the system to sustain 20% higher load, relative to traditional thread scheduling techniques. We implement and evaluate a suite of applications, including MICA, Raft and Set Algebra for document retrieval; and we demonstrate that the nanoPU can be used as a high performance, programmable alternative for one-sided RDMA operations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 67959401-e2af-4c97-89ec-85610a605deaCited by top-tier papers34
- Efficient Scheduling Policies for Microsecond-Scale TasksSarah McClure, Amy Ousterhout, Scott Shenker, Sylvia RatnasamyNSDI 2022 · 43 citations
- When Idling is Ideal: Optimizing Tail-Latency for Heavy-Tailed Datacenter Workloads with PerséphoneHenri Maxime Demoulin, Joshua Fried, Isaac Pedisich, Marios Kogias et al.SOSP 2021 · 39 citations
- A Cloud-Scale Characterization of Remote Procedure CallsKorakit Seemakhupt, Brent E. Stephens, Samira Manabi Khan, Sihang Liu et al.SOSP 2023 · 31 citations
- Cerebros: Evading the RPC Tax in DatacentersArash Pourhabibi Zarandi, Mark Sutherland, Alexandros Daglis, Babak FalsafiMICRO 2021 · 24 citations
- uBFT: Microsecond-Scale BFT using Disaggregated MemoryMarcos K. Aguilera, Naama Ben-David, Rachid Guerraoui, Antoine Murat et al.ASPLOS 2023 · 20 citations
Builds on6
- Disaggregating Persistent Memory and Controlling Them Remotely: An Exploration of Passive Disaggregated Key-Value StoresShin-Yeh Tsai, Yizhou Shan, Yiying ZhangUSENIX ATC 2020 · 159 citations
- Enabling Programmable Transport Protocols in High-Speed NICsMina Tahmasbi Arashloo, Alexey Lavrov, Manya Ghobadi, Jennifer Rexford et al.NSDI 2020 · 96 citations
- Reexamining Direct Cache Access to Optimize I/O Intensive Applications for Multi-hundred-gigabit NetworksAlireza Farshin, Amir Roozbeh, Gerald Q. Maguire Jr., Dejan KosticUSENIX ATC 2020 · 88 citations
- StRoM: smart remote memoryDavid Sidler, Zeke Wang, Monica Chiosa, Amit Kulkarni et al.EuroSys 2020 · 83 citations
- Optimus Prime: Accelerating Data Transformation in ServersArash Pourhabibi Zarandi, Siddharth Gupta, Hussein Kassir, Mark Sutherland et al.ASPLOS 2020 · 43 citations
Related papers
- ALTOCUMULUS: Scalable Scheduling for Nanosecond-Scale Remote Procedure CallsJiechen Zhao, Iris Uwizeyimana, Karthik Ganesan, Mark C. Jeffrey et al.MICRO 2022 · 11 citations
- RpcNIC: Enabling Efficient Datacenter RPC Offloading on PCIe-attached SmartNICsJie Zhang, Hongjing Huang, Xuzheng Chen, Xiang Li et al.HPCA 2025 · 6 citations
- SwCC: Software-Programmable and Per-Packet Congestion Control in RDMA EngineHongjing Huang, Jie Zhang, Xuzheng Chen, Ziyu Song et al.USENIX ATC 2025 · 4 citations
- Breaking Barriers in Atomic Scaling: A Hardware-Software-Collaborated Framework to Deconstruct RDMA AtomicGuangyang Deng, Qiangsheng Su, Zhirong Shen, Qing Wang et al.ISCA 2026
- Dagger: efficient and fast RPCs in cloud microservices with near-memory reconfigurable NICsNikita Lazarev, Shaojie Xiang, Neil Adit, Zhiru Zhang et al.ASPLOS 2021 · 54 citations
