AIO: An Abstraction for Performance Analysis Across Diverse Accelerator Architectures
Joseph Rogers, Taha Soliman, Magnus Jahre
摘要
Specialization is the key approach for continued performance growth beyond the end of Dennard scaling. Academics and industry are hence continuously proposing new accelerator architectures, including conventional Domain-Specific Accelerators (DSAs) and emerging Processing in Memory (PIM) accelerators. We are thus fast approaching an era in which earlystage accelerator analysis is critical for maintaining the productivity of software developers, system software designers, and computer architects — to ensure that they focus time-consuming implementation and optimization efforts on the most favorable class of accelerators for the problem at hand. Unfortunately, existing approaches fall short because they either adopt a level of abstraction that is too high — and therefore are unable to account for key performance phenomena — or too low — because they focus on details that do not generalize across diverse accelerators. Our Architecture-Independent Operation (AIO) abstraction addresses this issue by leveraging that accelerators typically focus on data-level parallelism, and an AIO is hence a key piece of algorithm-level data-parallel work that remains the same across diverse accelerators. To demonstrate that the AIO abstraction can be accurate and useful, we create the AccMe performance model which predicts kernel performance by estimating the number of clock cycles spent on compute, memory, and invocation overhead while accounting for overlap between compute and memory cycles as well as finite memory bandwidth. We demonstrate that AccMe can be accurate, i.e., it yields an average performance prediction error of across our diverse kernels and accelerators. This is a significant improvement over the average error of curve-fitted Roofline which provides the best-case accuracy of Roofline’s operational intensity abstraction. We further demonstrate that AccMe is useful through three case studies that illustrate (i) how developers can use AccMe for accelerator selection under uncertainty; (ii) how system software can use AccMe for scheduling — and thereby improve throughput by on average compared to Roofline-driven scheduling; and (iii) how computer architects can use AccMe for architectural exploration.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- HILP: Accounting for Workload-Level Parallelism in System-on-Chip Design Space ExplorationJoseph Rogers, Lieven Eeckhout, Magnus JahreHPCA 2025 · 被引用 4 次
- Neoscope: How Resilient Is My SoC to Workload Churn?Joseph Rogers, Lieven Eeckhout, Taha Soliman, Magnus JahreISCA 2025 · 被引用 2 次
相关 Paper
- The Configuration Wall: Characterization and Elimination of Accelerator Configuration OverheadJosse Van Delm, Anton Lydike, Joren Dumoulin, Jonas Crols 等ASPLOS 2026
- Pathfinding Future PIM Architectures by Demystifying a Commercial PIM TechnologyBongjoon Hyun, Taehun Kim, Dongjae Lee, Minsoo RhuHPCA 2024 · 被引用 62 次
- DCC: Data-Centric Compilation of Machine Learning Kernels for Processing-In-Memory ArchitecturesPeiming Yang, Sankeerth Durvasula, Ivan Fernandez, Mohammad Sadrosadati 等ISCA 2026 · 被引用 3 次
- Effectively Scheduling Computational Graphs of Deep Neural Networks toward Their Domain-Specific AcceleratorsJie Zhao, Siyuan Feng, Xiaoqiang Dan, Fei Liu 等OSDI 2023 · 被引用 9 次
- PIM-CCA: An Efficient PIM Architecture with Optimized Integration of Configurable Functional UnitsJeehyun Kim, Donghyeon Kim, Seokwon Kang, Bongjoon Hyun 等MICRO 2025 · 被引用 3 次
