Dynamic Detection of Inefficient Data Mapping Patterns in Heterogeneous OpenMP Applications
Luke Marzen, Junhyung Shim, Ali Jannesari
摘要
With the growing prevalence of heterogeneous computing, CPUs are increasingly being paired with accelerators to achieve new levels of performance and energy efficiency. However, data movement between devices remains a significant bottleneck, complicating application development. Existing performance tools require considerable programmer intervention to diagnose and locate data transfer inefficiencies. To address this, we propose dynamic analysis techniques to detect and profile inefficient data transfer and allocation patterns in heterogeneous applications. We implemented these techniques into OMPDataPerf, which provides detailed traces of problematic data mappings, source code attribution, and assessments of optimization potential in heterogeneous OpenMP applications. OMPDataPerf uses the OpenMP Tools Interface (OMPT) and incurs only a 5 % geometric‑mean runtime overhead.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- ValueExpert: exploring value patterns in GPU-accelerated applicationsKeren Zhou, Yueming Hao, John M. Mellor-Crummey, Xiaozhu Meng 等ASPLOS 2022 · 被引用 19 次
- DrGPUM: Guiding Memory Optimization for GPU-Accelerated ApplicationsMao Lin, Keren Zhou, Pengfei SuASPLOS 2023 · 被引用 13 次
- Static Generation of Efficient OpenMP Offload Data MappingsLuke Marzen, Akash Dutta, Ali JannesariSC 2024 · 被引用 4 次
相关 Paper
- CCAMP: an integrated translation and optimization framework for OpenACC and OpenMPJacob Lambert, Seyong Lee, Jeffrey S. Vetter, Allen D. MalonySC 2020 · 被引用 17 次
- OMPRacer: a scalable and precise static race detector for OpenMP programsBradley Swain, Yanze Li, Peiming Liu, Ignacio Laguna 等SC 2020 · 被引用 22 次
- GVARP: Detecting Performance Variance on Large-Scale Heterogeneous SystemsXin You, Zhibo Xuan, Hailong Yang, Zhongzhi Luan 等SC 2024 · 被引用 8 次
- Automated Mapping of Task-Based Programs onto Distributed and Heterogeneous MachinesThiago S. F. X. Teixeira, Alexandra Henzinger, Rohan Yadav, Alex AikenSC 2023 · 被引用 5 次
- Synergically Rebalancing Parallel Execution via DCT and Turbo BoostingSandro M. Marques, Thiarles S. Medeiros, Fábio Diniz Rossi, Marcelo Caggiani Luizelli 等DAC 2021 · 被引用 5 次
