Leveraging Chiplet-Locality for Efficient Memory Mapping in Multi-Chip Module GPUs
Junhyeok Park, Sungbin Jang, Osang Kwon, Yongho Lee, Seokin Hong
摘要
While the multi-chip module (MCM) design allows GPUs to scale compute and memory capabilities through multi-chip integration, it introduces memory system non-uniformity, particularly when a thread accesses resources in remote chiplets.In this work, we investigate how page size in memory mapping affects this nonuniformity.Large pages reduce address translation overhead by covering larger memory regions per TLB entry; however, they enforce coarse-grained data placement, which can lead to data misallocation across chiplets.In contrast, small pages allow for finer-grained placement, increasing the likelihood of mapping data to the chiplet most likely to access it.We observe that application performance is sensitive to page size, with the appropriate configuration depending on workload characteristics.This paper introduces CLAP which determines the suitable page size-specifically, how much data should be co-located within a single chiplet-for each application.We observe that GPU applications exhibit a distinct memory mapping pattern, in which specific groups of virtually adjacent pages are primarily accessed by the same chiplet with the group size tending to remain consistenta property referred to as chiplet-locality.Leveraging this insight, CLAP predicts groups of pages exhibit chiplet-locality and preorganizes them to contiguous physical frames within the chiplet most likely to access them.This organization forms regions that behave like large pages, as CLAP enables these page groups to be covered by a single merged TLB entry through deliberate virtualto-physical contiguity.As a result, CLAP delivers the benefits of large pages without compromising chiplet-level memory locality.Our evaluation shows that CLAP improves performance by up to 19.2% compared to previous paging schemes.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- WATOS: Efficient LLM Training Strategies and Architecture Co-Exploration for Wafer-Scale ChipHuizheng Wang, Zichuan Wang, Hongbin Wang, Jingxiang Hou 等HPCA 2026 · 被引用 2 次
- Revelator: Rapid Data Fetching Via System-Software-Guided Hash-Based Speculative Address TranslationKonstantinos Kanellopoulos, Konstantinos Sgouras, Harsh Songara, Andreas Kosmas Kakolyris 等ISCA 2026
- Understanding and Profiling the Accelerator Chiplet Network Using PingPointJunyeol Ryu, Ming Liu, Matthew D. SinclairSIGCOMM 2026
相关 Paper
- Designing Virtual Memory System of MCM GPUsPratheek B, Neha Jawalkar, Arkaprava BasuMICRO 2022 · 被引用 21 次
- Barre Chord: Efficient Virtual Memory Translation for Multi-Chip-Module GPUsYuan Feng, Seonjin Na, Hyesoon Kim, Hyeran JeonISCA 2024 · 被引用 20 次
- Locality-Centric Data and Threadblock Management for Massive GPUsMahmoud Khairy, Vadim Nikiforov, David W. Nellans, Timothy G. RogersMICRO 2020 · 被引用 38 次
- SAC: Sharing-Aware Caching in Multi-Chip GPUsShiqing Zhang, Mahmood Naderan-Tahan, Magnus Jahre, Lieven EeckhoutISCA 2023 · 被引用 18 次
- Batch-Aware Unified Memory Management in GPUs for Irregular WorkloadsHyojong Kim, Jaewoong Sim, Prasun Gera, Ramyad Hadidi 等ASPLOS 2020 · 被引用 89 次
