IOctopus: Outsmarting Nonuniform DMA
Igor Smolyar, Alex Markuze, Boris Pismenny, Haggai Eran, Gerd Zellweger, Austin Bolen, Liran Liss, Adam Morrison, Dan Tsafrir
摘要
In a multi-CPU server, memory modules are local to the CPU to which they are connected, forming a nonuniform memory access (NUMA) architecture. Because non-local accesses are slower than local accesses, the NUMA architecture might degrade application performance. Similar slowdowns occur when an I/O device issues nonuniform DMA (NUDMA) operations, as the device is connected to memory via a single CPU. NUDMA effects therefore degrade application performance similarly to NUMA effects.
We observe that the similarity is not inherent but rather a product of disregarding the intrinsic differences between I/O and CPU memory accesses. Whereas NUMA effects are inevitable, we show that NUDMA effects can and should be eliminated. We present IOctopus, a device architecture that makes NUDMA impossible by unifying multiple physical PCIe functions-one per CPU-in manner that makes them appear as one, both to the system software and externally to the server. IOctopus requires only a modest change to the device driver and firmware. We implement it on existing hardware and demonstrate that it improves throughput and latency by as much as 2.7× and 1.28×, respectively, while ridding developers from the need to combat (what appeared to be) an unavoidable type of overhead.
• Hardware → Communication hardware, interfaces and storage; • Software and its engineering → Operating systems; Input / output.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Reexamining Direct Cache Access to Optimize I/O Intensive Applications for Multi-hundred-gigabit NetworksAlireza Farshin, Amir Roozbeh, Gerald Q. Maguire Jr., Dejan KosticUSENIX ATC 2020 · 被引用 88 次
- Characterizing and Optimizing Remote Persistent Memory with RDMA and NVMXingda Wei, Xiating Xie, Rong Chen, Haibo Chen 等USENIX ATC 2021 · 被引用 48 次
- Nap: A Black-Box Approach to NUMA-Aware Persistent Memory IndexesQing Wang, Youyou Lu, Junru Li, Jiwu ShuOSDI 2021 · 被引用 46 次
- Don't Forget the I/O When Allocating Your LLCYifan Yuan, Mohammad Alian, Yipeng Wang, Ren Wang 等ISCA 2021 · 被引用 37 次
- Autonomous NIC offloadsBoris Pismenny, Haggai Eran, Aviad Yehezkel, Liran Liss 等ASPLOS 2021 · 被引用 32 次
相关 Paper
- PaCaR: Improved Buffered I/O Locality on NUMA Systems with Page Cache ReplicationJérôme Coquisart, Julien Sopena, Redha GouicemEuroSys 2026
- Optimizing Storage Performance with Calibrated InterruptsAmy Tai, Igor Smolyar, Michael Wei, Dan TsafrirOSDI 2021
- NUBA: Non-Uniform Bandwidth GPUsXia Zhao, Magnus Jahre, Yuhua Tang, Guangda Zhang 等ASPLOS 2023 · 被引用 17 次
- Efficient Remote Memory Ordering for Non-Coherent SystemsWei Siew Liew, Md Ashfaqur Rahaman, Adarsh Patil, Ryan Stutsman 等ASPLOS 2026
- DDS: DPU-optimized Disaggregated StorageQizhen Zhang, Philip A. Bernstein, Badrish Chandramouli, Jason Hu 等VLDB 2024 · 被引用 12 次
