IOctopus: Outsmarting Nonuniform DMA
Igor Smolyar, Alex Markuze, Boris Pismenny, Haggai Eran, Gerd Zellweger, Austin Bolen, Liran Liss, Adam Morrison, Dan Tsafrir
Abstract
In a multi-CPU server, memory modules are local to the CPU to which they are connected, forming a nonuniform memory access (NUMA) architecture. Because non-local accesses are slower than local accesses, the NUMA architecture might degrade application performance. Similar slowdowns occur when an I/O device issues nonuniform DMA (NUDMA) operations, as the device is connected to memory via a single CPU. NUDMA effects therefore degrade application performance similarly to NUMA effects.
We observe that the similarity is not inherent but rather a product of disregarding the intrinsic differences between I/O and CPU memory accesses. Whereas NUMA effects are inevitable, we show that NUDMA effects can and should be eliminated. We present IOctopus, a device architecture that makes NUDMA impossible by unifying multiple physical PCIe functions-one per CPU-in manner that makes them appear as one, both to the system software and externally to the server. IOctopus requires only a modest change to the device driver and firmware. We implement it on existing hardware and demonstrate that it improves throughput and latency by as much as 2.7× and 1.28×, respectively, while ridding developers from the need to combat (what appeared to be) an unavoidable type of overhead.
• Hardware → Communication hardware, interfaces and storage; • Software and its engineering → Operating systems; Input / output.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers13
- Reexamining Direct Cache Access to Optimize I/O Intensive Applications for Multi-hundred-gigabit NetworksAlireza Farshin, Amir Roozbeh, Gerald Q. Maguire Jr., Dejan KosticUSENIX ATC 2020 · 88 citations
- Characterizing and Optimizing Remote Persistent Memory with RDMA and NVMXingda Wei, Xiating Xie, Rong Chen, Haibo Chen et al.USENIX ATC 2021 · 48 citations
- Nap: A Black-Box Approach to NUMA-Aware Persistent Memory IndexesQing Wang, Youyou Lu, Junru Li, Jiwu ShuOSDI 2021 · 46 citations
- Don't Forget the I/O When Allocating Your LLCYifan Yuan, Mohammad Alian, Yipeng Wang, Ren Wang et al.ISCA 2021 · 37 citations
- Autonomous NIC offloadsBoris Pismenny, Haggai Eran, Aviad Yehezkel, Liran Liss et al.ASPLOS 2021 · 32 citations
Related papers
- PaCaR: Improved Buffered I/O Locality on NUMA Systems with Page Cache ReplicationJérôme Coquisart, Julien Sopena, Redha GouicemEuroSys 2026
- Optimizing Storage Performance with Calibrated InterruptsAmy Tai, Igor Smolyar, Michael Wei, Dan TsafrirOSDI 2021
- NUBA: Non-Uniform Bandwidth GPUsXia Zhao, Magnus Jahre, Yuhua Tang, Guangda Zhang et al.ASPLOS 2023 · 17 citations
- Efficient Remote Memory Ordering for Non-Coherent SystemsWei Siew Liew, Md Ashfaqur Rahaman, Adarsh Patil, Ryan Stutsman et al.ASPLOS 2026
- DDS: DPU-optimized Disaggregated StorageQizhen Zhang, Philip A. Bernstein, Badrish Chandramouli, Jason Hu et al.VLDB 2024 · 12 citations
