Toasty: Speeding Up Network I/O with Cache-Warm Buffers
Preeti, Nitish Bhat, Ashwin Kumar, Mythili Vutukuru
Abstract
Modern NICs DMA packets directly to the LLC using technologies like DDIO, reducing access latencies for networking applications. Prior work has observed several performance issues with DDIO when the working set of packet buffers does not fit into LLC. For example, the leaky DMA problem arises when incoming packets evict older packets that have not yet been processed by the application from LLC, causing them to be fetched again from main memory. While using a smaller pool of packet buffers that fits in cache is an obvious solution, this may result in the NIC running out of buffers to DMA packets into when a burst of packets arrives. This paper proposes Toasty, a system that mitigates this tradeoff between high throughput and resilience to packet loss that arises when sizing the network packet buffer pool. While prior work has proposed hardware-based solutions to this problem, Toasty is a software-only solution that can be deployed on commodity NIC hardware. Toasty manages the packet buffer pool as a LIFO stack instead of a FIFO queue, and adapts the number of buffers populated into the NIC hardware RX ring based on incoming packet load and application processing rate. Together, these changes enable Toasty to recirculate a small working set of cache-warm buffers in steady state, while falling back to a larger pool of buffers during traffic bursts. We implement Toasty over the AF_XDP kernel bypass framework, and our evaluation shows that Toasty improves network throughput for a variety of network functions by up to 78% over the default buffer pool implementation of AF_XDP. We also show that Toasty matches the performance of a small buffer pool that fits in cache, while being more resilient to traffic bursts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9599a383-aab6-4436-8e33-dd95a643c821Builds on9
- Caladan: Mitigating Interference at Microsecond TimescalesJoshua Fried, Zhenyuan Ruan, Amy Ousterhout, Adam BelayOSDI 2020 · 213 citations
- Reexamining Direct Cache Access to Optimize I/O Intensive Applications for Multi-hundred-gigabit NetworksAlireza Farshin, Amir Roozbeh, Gerald Q. Maguire Jr., Dejan KosticUSENIX ATC 2020 · 88 citations
- Don't Forget the I/O When Allocating Your LLCYifan Yuan, Mohammad Alian, Yipeng Wang, Ren Wang et al.ISCA 2021 · 37 citations
- The benefits of general-purpose on-NIC memoryBoris Pismenny, Liran Liss, Adam Morrison, Dan TsafrirASPLOS 2022 · 28 citations
- IDIO: Network-Driven, Inbound Network Data Orchestration on Server ProcessorsMohammad Alian, Siddharth Agarwal, Jongmin Shin, Neel Patel et al.MICRO 2022 · 21 citations
Related papers
- Patching up Network Data Leaks with SweeperMarina Vemmou, Albert Cho, Alexandros DaglisMICRO 2022 · 6 citations
- Disentangling the Dual Role of NIC Receive RingsBoris Pismenny, Adam Morrison, Dan TsafrirOSDI 2025 · 2 citations
- When DDIO Meets Page Coloring: Revisiting DDIO Performance with SepiaChangwoo Song, Sanghyun Kim, Jinhyeok Oh, Qizhe Cai et al.OSDI 2026
- CEIO: A Cache-Efficient Network I/O Architecture for NIC-CPU Data PathsBowen Liu, Xinyang Huang, Qijing Li, Zhuobin Huang et al.SIGCOMM 2025 · 10 citations
- GPU-Ether: GPU-native Packet I/O for GPU Applications on Commodity EthernetChangue Jung, Suhwan Kim, Ikjun Yeom, Honguk Woo et al.INFOCOM 2021 · 7 citations
