The Configuration Wall: Characterization and Elimination of Accelerator Configuration Overhead
Josse Van Delm, Anton Lydike, Joren Dumoulin, Jonas Crols, Xiaoling Yi, Ryan Antonio, Jackson Woodruff, Tobias Grosser, Marian Verhelst
Abstract
Contemporary compute platforms increasingly offload compute kernels from CPU to integrated hardware accelerators to reach maximum performance per Watt. Unfortunately, the time the CPU spends on setup control and synchronization has increased with growing accelerator complexity. For systems with complex accelerators, this means that performance can be configuration-bound. Faster accelerators are more severely impacted by this overlooked performance drop, which we call the configuration wall. Prior work evidences this wall and proposes ad-hoc solutions to reduce configuration overhead. However, these solutions are not universally applicable, nor do they offer comprehensive insights into the underlying causes of performance degradation. In this work, we first introduce a widely-applicable variant of the well-known roofline model to quantify when system performance is configuration-bound. To move systems out of the performance-bound region, we subsequently propose a domain-specific compiler abstraction and associated optimization passes. We implement the abstraction and passes in the MLIR compiler framework to run optimized binaries on open-source architectures to prove its effectiveness and generality. Experiments demonstrate a geomean performance boost of 2x on the open-source OpenGeMM system, by eliminating redundant configuration cycles and by automatically hiding the remaining configuration cycles. Our work provides key insights in how accelerator performance is affected by setup mechanisms, thereby facilitating automatic code generation for circumventing the configuration wall.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d116b5e8-9120-4a74-b854-0dfd48834403Cited by top-tier papers1
Ask how each one uses itBuilds on6
- Gemmini: Enabling Systematic Deep-Learning Architecture Evaluation via Full-Stack IntegrationHasan Genc, Seah Kim, Alon Amid, Ameer Haj-Ali et al.DAC 2021 · 325 citations
- Optimus Prime: Accelerating Data Transformation in ServersArash Pourhabibi Zarandi, Siddharth Gupta, Hussein Kassir, Mark Sutherland et al.ASPLOS 2020 · 43 citations
- A Hardware Accelerator for Protocol BuffersSagar Karandikar, Chris Leary, Chris Kennelly, Jerry Zhao et al.MICRO 2021 · 42 citations
- RACOD: algorithm/hardware co-design for mobile robot path planningMohammad Bakhshalipour, Seyed Borna Ehsani, Mohamad Qadri, Dominic Guri et al.ISCA 2022 · 21 citations
- Cohort: Software-Oriented Acceleration for Heterogeneous SoCsTianrui Wei, Nazerke Turtayeva, Marcelo Orenes-Vera, Omkar Lonkar et al.ASPLOS 2023 · 12 citations
Related papers
- AIO: An Abstraction for Performance Analysis Across Diverse Accelerator ArchitecturesJoseph Rogers, Taha Soliman, Magnus JahreISCA 2024 · 5 citations
- HIR: An MLIR-based Intermediate Representation for Hardware Accelerator DescriptionKingshuk Majumder, Uday BondhugulaASPLOS 2023 · 12 citations
- WASP: Exploiting GPU Pipeline Parallelism with Hardware-Accelerated Automatic Warp SpecializationNeal Clayton Crago, Sana Damani, Karthikeyan Sankaralingam, Stephen W. KecklerHPCA 2024 · 13 citations
- Squeezing Operator Performance Potential for the Ascend ArchitectureYuhang Zhou, Zhibin Wang, Guyue Liu, Shipeng Li et al.ASPLOS 2025 · 3 citations
- Compiler-Driven Simulation of Reconfigurable Hardware AcceleratorsZhijing Li, Yuwei Ye, Stephen Neuendorffer, Adrian SampsonHPCA 2022 · 4 citations
