Lune

ASPLOS2026顶会

TiNA: Tiered Network Buffer Architecture for Fast Networking in Chiplet-based CPUs

Siddharth Agarwal, Tianchen Wang, Jinghan Huang, Saksham Agarwal, Nam Sung Kim

2026年份
1被引次数

摘要

To manufacture a large CPU cost-effectively, the industry has begun exploiting emerging packaging technologies that integrate multiple chiplets-each comprising a subset of cores and/or memory and I/O subsystems-into a single package. However, such a CPU experiences longer memory access latency with more pronounced variance, especially when its cores in one chiplet access LLC slices 1 or DRAM controllers in other chiplets. This creates unique challenges in µs-scale networking, which is highly sensitive to memory access latency. In this work, we start by proposing exploiting a little-known mode, known as Sub-NUMA Clustering (SNC), in the latest chiplet-based CPUs. As it restricts receiving and processing packets to a particular chiplet unless explicitly specified otherwise, it offers shorter memory access latency and, consequently, lower networking latency than the default mode (non-SNC). Nonetheless, when receiving long bursts of packets 2 , SNC incurs higher networking latency than non-SNC, as it provides less LLC capacity for CPU cores processing the packets, making Direct Cache Access (DCA)a commonly used CPU feature to reduce memory access latency for packet processing-ineffective. To address this * Both authors contributed equally to this research. 1 A slice is a subset of the LLC directly connected with a CPU core. The CPU core can access the LLC slices of other CPU cores through interconnects but at longer latency than its own LLC slice. 2 A burst is defined as the period during which the NIC receives packets at line rate, and bursts lasting hundreds of microseconds frequently occur in datacenter network traffic [43]

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper7

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖