CHARM: Chiplet Heterogeneity-Aware Runtime Mapping System
Alessandro Fogli, Bo Zhao, Peter R. Pietzuch, Jana Giceva
Abstract
The growing disparity between CPU core counts and available memory bandwidth has intensified memory contention in servers. This particularly affects highly parallelizable applications, which must achieve efficient cache utilization to maintain performance as CPU core counts grow. Optimizing cache utilization, however, is complex for recent chipletbased CPUs, whose partitioned L3 caches lead to varying latencies and bandwidths, even within a single NUMA domain. Classical NUMA optimizations and task scheduling fail to address the performance issues of chiplet-based CPUs.
We describe Chiplet Heterogeneity Aware Runtime Mapping (CHARM), a new runtime system designed for chipletbased CPUs. CHARM combines chiplet-aware task scheduling heuristics, hardware-aware memory allocation, and fine-grained performance monitoring to optimize workload execution. It implements a lightweight concurrency model that combines user-level threading features, such as individual stacks, per-task scheduling, and state management, with coroutine-like behavior, allowing tasks to suspend and resume execution at defined points while efficiently managing task migration across chiplets. Our evaluation across diverse scenarios shows CHARM's effectiveness in optimizing the performance of memory-intensive parallel applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5af7f05e-1c97-4555-83b8-a2fefee40f0fCited by top-tier papers1
Ask how each one uses itBuilds on4
- Johnny Cache: the End of DRAM Cache Conflicts (in Tiered Main Memory Systems)Baptiste Lepers, Willy ZwaenepoelOSDI 2023 · 26 citations
- Put an Elephant into a Fridge: Optimizing Cache Efficiency for In-memory Key-value StoresKefei Wang, Jian Liu, Feng ChenVLDB 2020 · 23 citations
- On Querying Connected Components in Large Temporal GraphsHaoxuan Xie, Yixiang Fang, Yuyang Xia, Wensheng Luo et al.SIGMOD 2023 · 20 citations
- OLAP on Modern Chiplet-Based ProcessorsAlessandro Fogli, Bo Zhao, Peter R. Pietzuch, Maximilian Bandle et al.VLDB 2024 · 6 citations
Related papers
- TAPMM: A Traffic-Aware Page Mapping Method for Multi-level NUMA SystemsFengkun Dong, Guoqing Xiao, Haotian Wang, Yikun Hu et al.DAC 2024
- COMET: Communication and Memory Co-Design for Fine-Grained AI Inference in MCM AcceleratorsTaishu Sheng, Guangyu Sun, Dezun DongHPCA 2026
- A. Delegato: Locality-Aware Atomic Memory Operations on ChipletsVíctor Soria Pardos, Adrià Armejach, Tiago Mück, Darío Suárez Gracia et al.MICRO 2025 · 1 citation
- Scheduling Linux Threads under I/O Chiplet Wall Using cSwitchSeunghyun An, Joontaek Oh, Ming LiuSOSP 2026
- Merchandiser: Data Placement on Heterogeneous Memory for Task-Parallel HPC Applications with Load-Balance AwarenessZhen Xie, Jie Liu, Jiajia Li, Dong LiPPoPP 2023 · 18 citations
