ATX: Accelerator Task Extensions
Gerasimos Gerogiannis, Stijn Eyerman, Josep Torrellas, Wim Heirman
Abstract
Integrating accelerators in CPU multicores combines the benefits of accelerated computation with the flexibility and programmability of CPUs. CPU-integrated accelerators can be classified into In-Core Accelerators (ICAs), which typically reside inside the core's pipeline, and Out-of-Core Accelerators (OCAs), which are typically attached at the core's cache subsystem. Both designs have shortcomings: ICAs can be bottlenecked by the core's general-purpose memory access interface, while OCAs' core-accelerator interface can limit execution overlap and expose communication overheads. In this paper we introduce Near-Core Accelerators (NCAs), a new class of accelerators that aims to combine the advantages of ICAs and OCAs, and address their shortcomings. Specifically, NCAs have their own read interface to the memory system, eliminating ICAs' bottleneck. At the same time, NCAs can be invoked speculatively and out-of-order, enabling both coreaccelerator execution overlap and low-overhead core-accelerator communication. We also propose the Accelerator Task Extensions (ATX), a set of instructions and hardware extensions to support NCAs. With ATX instructions, CPU cores can speculatively invoke a diverse range of NCAs. ATX includes the Unified Transfer Engine (UTE), a programmable hardware module that efficiently supplies data to NCAs and virtualizes the interface between the CPU core and the NCAs. The UTE interfaces the CPU core with multiple NCAs and the cache subsystem, fetching and prefetching accelerator data, and scheduling tasks to potentially multiple NCAs transparently to the CPU core. We evaluate ATX NCAs with a variety of important kernels from machine learning and scientific computing. We show that ATX NCAs accelerate these kernels by over various CPUintegrated accelerator alternatives.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get d3c73a6c-49ec-4434-a3a5-e887dbab970dRelated papers
- A Tensor Marshaling Unit for Sparse Tensor Algebra on General-Purpose ProcessorsMarco Siracusa, Víctor Soria Pardos, Francesco Sgherzi, Joshua Randall et al.MICRO 2023 · 11 citations
- ARCANE: Adaptive RISC-V Cache Architecture for Near-memory ExtensionsVincenzo Petrolo, Flavia Guella, Michele Caon, Pasquale Davide Schiavone et al.DAC 2025 · 1 citation
- BlueFace: Integrating an Accelerator into the Core's Pipeline through Algorithm-Interface Co-Design for Real-Time SoCsZhe Jiang, Nathan Fisher, Nan Guan, Zheng DongDAC 2023 · 1 citation
- NOVIA: A Framework for Discovering Non-Conventional Inline AcceleratorsDavid Trilla, John-David Wellman, Alper Buyuktosunoglu, Pradip BoseMICRO 2021 · 14 citations
- An architecture interface and offload model for low-overhead, near-data, distributed acceleratorsSaambhavi Baskaran, Mahmut Taylan Kandemir, Jack SampsonMICRO 2022 · 13 citations
