Merchandiser: Data Placement on Heterogeneous Memory for Task-Parallel HPC Applications with Load-Balance Awareness
Zhen Xie, Jie Liu, Jiajia Li, Dong Li
摘要
The emergence of heterogeneous memory (HM) provides a cost-effective and high-performance solution to memory-consuming HPC applications. Deciding the placement of data objects on HM is critical for high performance. We reveal a performance problem related to data placement on HM. The problem is manifested as load imbalance among tasks in task-parallel HPC applications. The root of the problem comes from being unaware of parallel-task semantics and an incorrect assumption that bringing frequently accessed pages to fast memory always leads to better performance. To address this problem, we introduce a load balance-aware page management system, named Merchandiser. Merchandiser introduces task semantics during memory profiling, rather than being application-agnostic. Using the limited task semantics, Merchandiser effectively sets up coordination among tasks on the usage of HM to finish all tasks fast instead of only considering any individual task. Merchandiser is highly automated to enable high usability. Evaluating with memory-consuming HPC applications, we show that Merchandiser reduces load imbalance and leads to an average of 17.1% and 15.4% (up to 26.0% and 23.2%) performance improvement, compared with a hardware-based solution and an industry-quality software-based solution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- FlexMem: Adaptive Page Profiling and Migration for Tiered MemoryDong Xu, Junhee Ryu, Kwangsik Shin, Pengfei Su 等USENIX ATC 2024 · 被引用 41 次
- Scalable Tuning of (OpenMP) GPU Applications via Kernel Record and ReplayKonstantinos Parasyris, Giorgis Georgakoudis, Esteban Rangel, Ignacio Laguna 等SC 2023 · 被引用 15 次
- Machine Learning-Guided Memory Optimization for DLRM Inference on Tiered MemoryJie Ren, Bin Ma, Shuangyan Yang, Benjamin Francis 等HPCA 2025 · 被引用 6 次
- cMPI: Using CXL Memory Sharing for MPI One-Sided and Two-Sided Inter-Node CommunicationsXi Wang, Bin Ma, Jongryool Kim, Byungil Koh 等SC 2025 · 被引用 5 次
- Getting a Handle on Unmanaged MemoryNick Wanninger, Tommy McMichen, Simone Campanoni, Peter A. DindaASPLOS 2024 · 被引用 3 次
它引用的顶会 Paper11
- ZeRO-Offload: Democratizing Billion-Scale Model TrainingJie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase 等USENIX ATC 2021 · 被引用 657 次
- An Empirical Guide to the Behavior and Use of Scalable Persistent MemoryJian Yang, Juno Kim, Morteza Hoseinzadeh, Joseph Izraelevitz 等FAST 2020 · 被引用 470 次
- HM-ANN: Efficient Billion-Point Nearest Neighbor Search on Heterogeneous MemoryJie Ren, Minjia Zhang, Dong LiNeurIPS 2020 · 被引用 136 次
- Exploring the Design Space of Page Management for Multi-Tiered Memory SystemsJonghyeon Kim, Wonkyo Choe, Jeongseob AhnUSENIX ATC 2021 · 被引用 108 次
- Single Machine Graph Analytics on Massive Datasets Using Intel Optane DC Persistent MemoryGurbinder Gill, Roshan Dathathri, Loc Hoang, Ramesh Peri 等VLDB 2020 · 被引用 82 次
相关 Paper
- A Quantitative Approach for Adopting Disaggregated Memory in HPC SystemsJacob Wahlgren, Gabin Schieffer, Maya B. Gokhale, Ivy PengSC 2023 · 被引用 14 次
- Sentinel: Efficient Tensor Migration and Allocation on Heterogeneous Memory Systems for Deep LearningJie Ren, Jiaolin Luo, Kai Wu, Minjia Zhang 等HPCA 2021 · 被引用 62 次
- TAPMM: A Traffic-Aware Page Mapping Method for Multi-level NUMA SystemsFengkun Dong, Guoqing Xiao, Haotian Wang, Yikun Hu 等DAC 2024
- TPP: Transparent Page Placement for CXL-Enabled Tiered-MemoryHasan Al Maruf, Hao Wang, Abhishek Dhanotia, Johannes Weiner 等ASPLOS 2023 · 被引用 255 次
- KLOCs: kernel-level object contexts for heterogeneous memory systemsSudarsun Kannan, Yujie Ren, Abhishek BhattacharjeeASPLOS 2021 · 被引用 22 次
