HyperPlane: A Scalable Low-Latency Notification Accelerator for Software Data Planes
Amirhossein Mirhosseini, Hossein Golestani, Thomas F. Wenisch
Abstract
I/O software stacks have evolved rapidly due to the growing speed of I/O devices-including network adapters, storage devices, and accelerators-and the emergence of microservice-based programming models. Datacenters rely on fast, efficient Software Data Planes (SDPs), which orchestrate data transfer between applications and I/O devices. Modern data planes are user-level software stacks, wherein cores spin-poll a large number of queues to avoid the attendant overheads of kernel-based I/O. Cores often poll empty queues before finding work in non-empty ones. Interrogating empty queues hurts peak throughput, tail latency, and energy efficiency as it often entails fruitless cache misses. In this work, we propose HyperPlane, an efficient accelerator for the notification mechanism of SDPs.
The key features of HyperPlane are (1) avoiding iteration over empty I/O queues, unlike software-only designs, resulting in queue scalability, (2) halting execution when I/O queues are idle, leading to work proportionality and energy efficiency, and (3) efficiently sharing queues across cores to enjoy strong theoretical properties of scale-up queuing. HyperPlane is realized through a hardware subsystem associated with a familiar programming model. HyperPlane's microarchitecture consists of a monitoring set that watches for work arrival from I/O, and a ready set, which tracks ready queues and distributes work to cores based on various service policies and priority levels. We show that HyperPlane improves peak throughput by 4.1× and tail latency by 16.4× compared to a state-of-the-art SDP.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7e5a595f-e46f-4312-8254-aa0dcaa35ab0Cited by top-tier papers8
- The benefits of general-purpose on-NIC memoryBoris Pismenny, Liran Liss, Adam Morrison, Dan TsafrirASPLOS 2022 · 28 citations
- SmartDS: Middle-Tier-centric SmartNIC Enabling Application-aware Message Split for Disaggregated Block StorageJie Zhang, Hongjing Huang, Lingjun Zhu, Shu Ma et al.ISCA 2023 · 15 citations
- Cohort: Software-Oriented Acceleration for Heterogeneous SoCsTianrui Wei, Nazerke Turtayeva, Marcelo Orenes-Vera, Omkar Lonkar et al.ASPLOS 2023 · 12 citations
- Extended User Interrupts (xUI): Fast and Flexible Notification without PollingBerk Aydogmus, Linsong Guo, Danial Zuberi, Tal Garfinkel et al.ASPLOS 2025 · 8 citations
- HardHarvest: Hardware-Supported Core Harvesting for MicroservicesJovan Stojkovic, Chunao Liu, Muhammad Shahbaz, Josep TorrellasISCA 2025 · 4 citations
Builds on4
- Accelerometer: Understanding Acceleration Opportunities for Data Center Overheads at HyperscaleAkshitha Sriraman, Abhishek DhanotiaASPLOS 2020 · 78 citations
- LeapIO: Efficient and Portable Virtual NVMe Storage on ARM SoCsHuaicheng Li, Mingzhe Hao, Stanko Novakovic, Vaibhav Gogte et al.ASPLOS 2020 · 58 citations
- Optimus Prime: Accelerating Data Transformation in ServersArash Pourhabibi Zarandi, Siddharth Gupta, Hussein Kassir, Mark Sutherland et al.ASPLOS 2020 · 43 citations
- Q-Zilla: A Scheduling Framework and Core Microarchitecture for Tail-Tolerant MicroservicesAmirhossein Mirhosseini, Brendan L. West, Geoffrey W. Blake, Thomas F. WenischHPCA 2020 · 30 citations
Related papers
- PAIO: General, Portable I/O Optimizations With Minor Application ModificationsRicardo Macedo, Yusuke Tanimura, Jason Haga, Vijay Chidambaram et al.FAST 2022
- Sassy: SmartNIC-Assisted Notification Delivery for μs-Scale RDMA WorkloadsHamed Seyedroudbari, Alexandros DaglisHPCA 2026
- LibPreemptible: Enabling Fast, Adaptive, and Hardware-Assisted User-Space SchedulingYueying Li, Nikita Lazarev, David Koufaty, Tenny Yin et al.HPCA 2024 · 14 citations
- Hyperion: Co-Optimizing SSD Access and GPU Computation for Cost-Efficient GNN TrainingJie Sun, Mo Sun, Zheng Zhang, Zuocheng Shi et al.ICDE 2025 · 3 citations
- When Idling is Ideal: Optimizing Tail-Latency for Heavy-Tailed Datacenter Workloads with PerséphoneHenri Maxime Demoulin, Joshua Fried, Isaac Pedisich, Marios Kogias et al.SOSP 2021 · 39 citations
