CPElide: Efficient Multi-Chiplet GPU Implicit Synchronization
Preyesh Dalmia, Rajesh Shashi Kumar, Matthew D. Sinclair
摘要
Chiplets are transforming computer system designs, allowing system designers to combine heterogeneous computing resources at unprecedented scales. Breaking larger, mono-lithic chips into smaller, connected chip lets helps performance continue scaling, avoids die size limitations, improves yield, and reduces design and integration costs. However, chip let-based designs introduce an additional level of hierarchy, which causes indirection and non-uniformity. This clashes with typ-ical heterogeneous systems: unlike CPU-based multi-chiplet systems, heterogeneous systems do not have significant OS support or complex coherence protocols to mitigate the impact of this indirection. Thus, exploiting locality across application phases is harder in multi-chiplet heterogeneous systems. We propose CPElide, which utilizes information already avail-able in heterogeneous systems' embedded microprocessor (the command processor) to track inter-chiplet data dependencies and aggressively perform implicit synchronization only when necessary, instead of conservatively like the state-of-the-art HMG. Across 24 workloads CPElide improves average performance (13%, 19%), energy (14%, 11 %), and network traffic (14%,17%), respectively, over current approaches and HMG.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper12
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein 等ASPLOS 2024 · 被引用 693 次
- Accel-Sim: An Extensible Simulation Framework for Validated GPU ModelingMahmoud Khairy, Zhesheng Shen, Tor M. Aamodt, Timothy G. RogersISCA 2020 · 被引用 366 次
- RecNMP: Accelerating Personalized Recommendation with Near-Memory ProcessingLiu Ke, Udit Gupta, Benjamin Youngjae Cho, David Brooks 等ISCA 2020 · 被引用 235 次
- PREMA: A Predictive Multi-Task Scheduling Algorithm For Preemptible Neural Processing UnitsYujeong Choi, Minsoo RhuHPCA 2020 · 被引用 150 次
- Locality-Centric Data and Threadblock Management for Massive GPUsMahmoud Khairy, Vadim Nikiforov, David W. Nellans, Timothy G. RogersMICRO 2020 · 被引用 38 次
相关 Paper
- Heterogeneous Die-to-Die Interfaces: Enabling More Flexible Chiplet Interconnection SystemsYinxiao Feng, Dong Xiang, Kaisheng MaMICRO 2023 · 被引用 13 次
- OLAP on Modern Chiplet-Based ProcessorsAlessandro Fogli, Bo Zhao, Peter R. Pietzuch, Maximilian Bandle 等VLDB 2024 · 被引用 6 次
- PhaseWeave: Phase-Aware Execution on Heterogeneous Chiplet Architectures for DatacentersJoshua Kim, Chaojie Zhang, Íñigo Goiri, Christopher J. Rossbach 等ISCA 2026 · 被引用 1 次
- Leveraging Chiplet-Locality for Efficient Memory Mapping in Multi-Chip Module GPUsJunhyeok Park, Sungbin Jang, Osang Kwon, Yongho Lee 等MICRO 2025 · 被引用 7 次
- ECO-CHIP: Estimation of Carbon Footprint of Chiplet-based Architectures for Sustainable VLSIChetan Choppali Sudarshan, Nikhil Matkar, Sarma B. K. Vrudhula, Sachin S. Sapatnekar 等HPCA 2024 · 被引用 54 次
