TiNA: Tiered Network Buffer Architecture for Fast Networking in Chiplet-based CPUs
Siddharth Agarwal, Tianchen Wang, Jinghan Huang, Saksham Agarwal, Nam Sung Kim
摘要
To manufacture a large CPU cost-effectively, the industry has begun exploiting emerging packaging technologies that integrate multiple chiplets-each comprising a subset of cores and/or memory and I/O subsystems-into a single package. However, such a CPU experiences longer memory access latency with more pronounced variance, especially when its cores in one chiplet access LLC slices 1 or DRAM controllers in other chiplets. This creates unique challenges in µs-scale networking, which is highly sensitive to memory access latency. In this work, we start by proposing exploiting a little-known mode, known as Sub-NUMA Clustering (SNC), in the latest chiplet-based CPUs. As it restricts receiving and processing packets to a particular chiplet unless explicitly specified otherwise, it offers shorter memory access latency and, consequently, lower networking latency than the default mode (non-SNC). Nonetheless, when receiving long bursts of packets 2 , SNC incurs higher networking latency than non-SNC, as it provides less LLC capacity for CPU cores processing the packets, making Direct Cache Access (DCA)a commonly used CPU feature to reduce memory access latency for packet processing-ineffective. To address this * Both authors contributed equally to this research. 1 A slice is a subset of the LLC directly connected with a CPU core. The CPU core can access the LLC slices of other CPU cores through interconnects but at longer latency than its own LLC slice. 2 A burst is defined as the period during which the NIC receives packets at line rate, and bursts lasting hundreds of microseconds frequently occur in datacenter network traffic [43]
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Attack Directories, Not Caches: Side Channel Attacks in a Non-Inclusive WorldMengjia Yan, Read Sprabery, Bhargava Gopireddy, Christopher W. Fletcher 等S&P 2019 · 被引用 201 次
- Reexamining Direct Cache Access to Optimize I/O Intensive Applications for Multi-hundred-gigabit NetworksAlireza Farshin, Amir Roozbeh, Gerald Q. Maguire Jr., Dejan KosticUSENIX ATC 2020 · 被引用 88 次
- Host Congestion ControlSaksham Agarwal, Arvind Krishnamurthy, Rachit AgarwalSIGCOMM 2023 · 被引用 47 次
- Understanding the Host NetworkMidhul Vuppalapati, Saksham Agarwal, Henry Schuh, Baris Kasikci 等SIGCOMM 2024 · 被引用 25 次
- IDIO: Network-Driven, Inbound Network Data Orchestration on Server ProcessorsMohammad Alian, Siddharth Agarwal, Jongmin Shin, Neel Patel 等MICRO 2022 · 被引用 21 次
相关 Paper
- Disentangling the Dual Role of NIC Receive RingsBoris Pismenny, Adam Morrison, Dan TsafrirOSDI 2025 · 被引用 2 次
- ShRing: Networking with Shared Receive RingsBoris Pismenny, Adam Morrison, Dan TsafrirOSDI 2023 · 被引用 13 次
- A4: Microarchitecture-Aware LLC Management for Datacenter Servers with Emerging I/O DevicesHaneul Park, Jiaqi Lou, Sangjin Lee, Yifan Yuan 等ISCA 2025 · 被引用 2 次
- TAPMM: A Traffic-Aware Page Mapping Method for Multi-level NUMA SystemsFengkun Dong, Guoqing Xiao, Haotian Wang, Yikun Hu 等DAC 2024
- : Practical Cache Attacks from the NetworkMichael Kurth, Ben Gras, Dennis Andriesse, Cristiano Giuffrida 等S&P 2020 · 被引用 78 次
