CHARM: Chiplet Heterogeneity-Aware Runtime Mapping System
Alessandro Fogli, Bo Zhao, Peter R. Pietzuch, Jana Giceva
摘要
The growing disparity between CPU core counts and available memory bandwidth has intensified memory contention in servers. This particularly affects highly parallelizable applications, which must achieve efficient cache utilization to maintain performance as CPU core counts grow. Optimizing cache utilization, however, is complex for recent chipletbased CPUs, whose partitioned L3 caches lead to varying latencies and bandwidths, even within a single NUMA domain. Classical NUMA optimizations and task scheduling fail to address the performance issues of chiplet-based CPUs.
We describe Chiplet Heterogeneity Aware Runtime Mapping (CHARM), a new runtime system designed for chipletbased CPUs. CHARM combines chiplet-aware task scheduling heuristics, hardware-aware memory allocation, and fine-grained performance monitoring to optimize workload execution. It implements a lightweight concurrency model that combines user-level threading features, such as individual stacks, per-task scheduling, and state management, with coroutine-like behavior, allowing tasks to suspend and resume execution at defined points while efficiently managing task migration across chiplets. Our evaluation across diverse scenarios shows CHARM's effectiveness in optimizing the performance of memory-intensive parallel applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper4
- Johnny Cache: the End of DRAM Cache Conflicts (in Tiered Main Memory Systems)Baptiste Lepers, Willy ZwaenepoelOSDI 2023 · 被引用 26 次
- Put an Elephant into a Fridge: Optimizing Cache Efficiency for In-memory Key-value StoresKefei Wang, Jian Liu, Feng ChenVLDB 2020 · 被引用 23 次
- On Querying Connected Components in Large Temporal GraphsHaoxuan Xie, Yixiang Fang, Yuyang Xia, Wensheng Luo 等SIGMOD 2023 · 被引用 20 次
- OLAP on Modern Chiplet-Based ProcessorsAlessandro Fogli, Bo Zhao, Peter R. Pietzuch, Maximilian Bandle 等VLDB 2024 · 被引用 6 次
相关 Paper
- TAPMM: A Traffic-Aware Page Mapping Method for Multi-level NUMA SystemsFengkun Dong, Guoqing Xiao, Haotian Wang, Yikun Hu 等DAC 2024
- COMET: Communication and Memory Co-Design for Fine-Grained AI Inference in MCM AcceleratorsTaishu Sheng, Guangyu Sun, Dezun DongHPCA 2026
- A. Delegato: Locality-Aware Atomic Memory Operations on ChipletsVíctor Soria Pardos, Adrià Armejach, Tiago Mück, Darío Suárez Gracia 等MICRO 2025 · 被引用 1 次
- Scheduling Linux Threads under I/O Chiplet Wall Using cSwitchSeunghyun An, Joontaek Oh, Ming LiuSOSP 2026
- Merchandiser: Data Placement on Heterogeneous Memory for Task-Parallel HPC Applications with Load-Balance AwarenessZhen Xie, Jie Liu, Jiajia Li, Dong LiPPoPP 2023 · 被引用 18 次
