TiNA: Tiered Network Buffer Architecture for Fast Networking in Chiplet-based CPUs
Siddharth Agarwal, Tianchen Wang, Jinghan Huang, Saksham Agarwal, Nam Sung Kim
Abstract
To manufacture a large CPU cost-effectively, the industry has begun exploiting emerging packaging technologies that integrate multiple chiplets-each comprising a subset of cores and/or memory and I/O subsystems-into a single package. However, such a CPU experiences longer memory access latency with more pronounced variance, especially when its cores in one chiplet access LLC slices 1 or DRAM controllers in other chiplets. This creates unique challenges in µs-scale networking, which is highly sensitive to memory access latency. In this work, we start by proposing exploiting a little-known mode, known as Sub-NUMA Clustering (SNC), in the latest chiplet-based CPUs. As it restricts receiving and processing packets to a particular chiplet unless explicitly specified otherwise, it offers shorter memory access latency and, consequently, lower networking latency than the default mode (non-SNC). Nonetheless, when receiving long bursts of packets 2 , SNC incurs higher networking latency than non-SNC, as it provides less LLC capacity for CPU cores processing the packets, making Direct Cache Access (DCA)a commonly used CPU feature to reduce memory access latency for packet processing-ineffective. To address this * Both authors contributed equally to this research. 1 A slice is a subset of the LLC directly connected with a CPU core. The CPU core can access the LLC slices of other CPU cores through interconnects but at longer latency than its own LLC slice. 2 A burst is defined as the period during which the NIC receives packets at line rate, and bursts lasting hundreds of microseconds frequently occur in datacenter network traffic [43]
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 94c3da74-3465-4e70-aaee-5626ad396ea6Builds on7
- Attack Directories, Not Caches: Side Channel Attacks in a Non-Inclusive WorldMengjia Yan, Read Sprabery, Bhargava Gopireddy, Christopher W. Fletcher et al.S&P 2019 · 201 citations
- Reexamining Direct Cache Access to Optimize I/O Intensive Applications for Multi-hundred-gigabit NetworksAlireza Farshin, Amir Roozbeh, Gerald Q. Maguire Jr., Dejan KosticUSENIX ATC 2020 · 88 citations
- Host Congestion ControlSaksham Agarwal, Arvind Krishnamurthy, Rachit AgarwalSIGCOMM 2023 · 47 citations
- Understanding the Host NetworkMidhul Vuppalapati, Saksham Agarwal, Henry Schuh, Baris Kasikci et al.SIGCOMM 2024 · 25 citations
- IDIO: Network-Driven, Inbound Network Data Orchestration on Server ProcessorsMohammad Alian, Siddharth Agarwal, Jongmin Shin, Neel Patel et al.MICRO 2022 · 21 citations
Related papers
- Disentangling the Dual Role of NIC Receive RingsBoris Pismenny, Adam Morrison, Dan TsafrirOSDI 2025 · 2 citations
- ShRing: Networking with Shared Receive RingsBoris Pismenny, Adam Morrison, Dan TsafrirOSDI 2023 · 13 citations
- A4: Microarchitecture-Aware LLC Management for Datacenter Servers with Emerging I/O DevicesHaneul Park, Jiaqi Lou, Sangjin Lee, Yifan Yuan et al.ISCA 2025 · 2 citations
- TAPMM: A Traffic-Aware Page Mapping Method for Multi-level NUMA SystemsFengkun Dong, Guoqing Xiao, Haotian Wang, Yikun Hu et al.DAC 2024
- : Practical Cache Attacks from the NetworkMichael Kurth, Ben Gras, Dennis Andriesse, Cristiano Giuffrida et al.S&P 2020 · 78 citations
