Understanding and Profiling the Accelerator Chiplet Network Using PingPoint
Junyeol Ryu, Ming Liu, Matthew D. Sinclair
Abstract
Emerging chiplet-based accelerators introduce a new class of intrahost networks-the Accelerator Chiplet Network (ACN)-that links compute chiplets, IO chiplets, and memory modules and increasingly governs application performance. Yet ACN behavior remains largely opaque: existing tools overlook on-package communication and instead attribute overheads to compute or memory subsystems, while ACN-induced latency, bandwidth heterogeneity, and congestion are hard to observe due to proprietary microarchitectures, tight coupling with the execution pipeline, and complex mappings between application activity and hardware.
To overcome this challenge, we build an ACN characterization framework that enables fine-grained, topology-aware probing of paths and links. We then use it to uncover fundamental ACN performance properties on multi-chiplet GPUs. Guided by these insights, we design PingPoint, a lightweight utility for ACN-native profiling. Our key insight is that modeling the ACN as a logical, hose-based graph with queueing abstractions, combined with in-situ software probing, makes systematic dissection of the otherwise opaque ACN possible. It injects latency and bandwidth probes while co-executing target kernels, captures cycle-level link-and path-granular distributions, and applies differential attribution to localize congestion to individual ACN links. Across diverse workloads and hardware, it exposes hidden bottlenecks, guides kernel placement and traffic shaping, quantifies the performance impact of ACN contention, and enables practical optimization with marginal overhead.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a158e53d-8bb6-49ea-9550-b31506554414Builds on35
- Mixture-of-Experts with Expert Choice RoutingYanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du et al.NeurIPS 2022 · 933 citations
- Accel-Sim: An Extensible Simulation Framework for Validated GPU ModelingMahmoud Khairy, Zhesheng Shen, Tor M. Aamodt, Timothy G. RogersISCA 2020 · 366 citations
- PINT: Probabilistic In-band Network TelemetryRan Ben Basat, Sivaramakrishnan Ramanathan, Yuliang Li, Gianni Antichi et al.SIGCOMM 2020 · 268 citations
- Understanding host network stack overheadsQizhe Cai, Shubham Chaudhary, Midhul Vuppalapati, Jaehyun Hwang et al.SIGCOMM 2021 · 150 citations
- Programmable Calendar Queues for High-speed Packet SchedulingNaveen Kr. Sharma, Chenxingyu Zhao, Ming Liu, Pravein G. Kannan et al.NSDI 2020 · 119 citations
Related papers
- Scheduling Linux Threads under I/O Chiplet Wall Using cSwitchSeunghyun An, Joontaek Oh, Ming LiuSOSP 2026
- COMET: Communication and Memory Co-Design for Fine-Grained AI Inference in MCM AcceleratorsTaishu Sheng, Guangyu Sun, Dezun DongHPCA 2026
- Evaluating Chiplet-based Large-Scale Interconnection Networks via Cycle-Accurate Packet-Parallel SimulationYinxiao Feng, Yuchen Wei, Dong Xiang, Kaisheng MaUSENIX ATC 2024 · 21 citations
- A Versatile and Flexible Chiplet-based System Design for Heterogeneous Manycore ArchitecturesHao Zheng, Ke Wang, Ahmed LouriDAC 2020 · 38 citations
- Kite: A Family of Heterogeneous Interposer Topologies Enabled via Accurate Interconnect ModelingSrikant Bharadwaj, Jieming Yin, Bradford M. Beckmann, Tushar KrishnaDAC 2020 · 94 citations
