CPElide: Efficient Multi-Chiplet GPU Implicit Synchronization
Preyesh Dalmia, Rajesh Shashi Kumar, Matthew D. Sinclair
Abstract
Chiplets are transforming computer system designs, allowing system designers to combine heterogeneous computing resources at unprecedented scales. Breaking larger, mono-lithic chips into smaller, connected chip lets helps performance continue scaling, avoids die size limitations, improves yield, and reduces design and integration costs. However, chip let-based designs introduce an additional level of hierarchy, which causes indirection and non-uniformity. This clashes with typ-ical heterogeneous systems: unlike CPU-based multi-chiplet systems, heterogeneous systems do not have significant OS support or complex coherence protocols to mitigate the impact of this indirection. Thus, exploiting locality across application phases is harder in multi-chiplet heterogeneous systems. We propose CPElide, which utilizes information already avail-able in heterogeneous systems' embedded microprocessor (the command processor) to track inter-chiplet data dependencies and aggressively perform implicit synchronization only when necessary, instead of conservatively like the state-of-the-art HMG. Across 24 workloads CPElide improves average performance (13%, 19%), energy (14%, 11 %), and network traffic (14%,17%), respectively, over current approaches and HMG.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5c55ab48-bbc8-4fda-a831-470259f90239Cited by top-tier papers1
Ask how each one uses itBuilds on12
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein et al.ASPLOS 2024 · 693 citations
- Accel-Sim: An Extensible Simulation Framework for Validated GPU ModelingMahmoud Khairy, Zhesheng Shen, Tor M. Aamodt, Timothy G. RogersISCA 2020 · 366 citations
- RecNMP: Accelerating Personalized Recommendation with Near-Memory ProcessingLiu Ke, Udit Gupta, Benjamin Youngjae Cho, David Brooks et al.ISCA 2020 · 235 citations
- PREMA: A Predictive Multi-Task Scheduling Algorithm For Preemptible Neural Processing UnitsYujeong Choi, Minsoo RhuHPCA 2020 · 150 citations
- Locality-Centric Data and Threadblock Management for Massive GPUsMahmoud Khairy, Vadim Nikiforov, David W. Nellans, Timothy G. RogersMICRO 2020 · 38 citations
Related papers
- Heterogeneous Die-to-Die Interfaces: Enabling More Flexible Chiplet Interconnection SystemsYinxiao Feng, Dong Xiang, Kaisheng MaMICRO 2023 · 13 citations
- OLAP on Modern Chiplet-Based ProcessorsAlessandro Fogli, Bo Zhao, Peter R. Pietzuch, Maximilian Bandle et al.VLDB 2024 · 6 citations
- PhaseWeave: Phase-Aware Execution on Heterogeneous Chiplet Architectures for DatacentersJoshua Kim, Chaojie Zhang, Íñigo Goiri, Christopher J. Rossbach et al.ISCA 2026 · 1 citation
- Leveraging Chiplet-Locality for Efficient Memory Mapping in Multi-Chip Module GPUsJunhyeok Park, Sungbin Jang, Osang Kwon, Yongho Lee et al.MICRO 2025 · 7 citations
- ECO-CHIP: Estimation of Carbon Footprint of Chiplet-based Architectures for Sustainable VLSIChetan Choppali Sudarshan, Nikhil Matkar, Sarma B. K. Vrudhula, Sachin S. Sapatnekar et al.HPCA 2024 · 54 citations
