An architecture interface and offload model for low-overhead, near-data, distributed accelerators
Saambhavi Baskaran, Mahmut Taylan Kandemir, Jack Sampson
摘要
The performance and energy costs of coordinating and performing data movement have led to proposals adding compute units and/or specialized access units to the memory hierarchy. However, current on-chip offload models are restricted to fixed compute and access pattern types, which limits software-driven optimizations and the applicability of such an offload interface to heterogeneous accelerator resources. This paper presents a computation offload interface for multi-core systems augmented with distributed on-chip accelerators. With energy-efficiency as the primary goal, we define mechanisms to identify offload partitioning, create a low-overhead execution model to sequence these fine-grained operations, and evaluate a set of workloads to identify the complexity needed to achieve distributed near-data execution. We demonstrate that our model and interface, combining features of dataflow in parallel with near-data processing engines, can be profitably applied to memory hierarchies augmented with either specialized compute substrates or lightweight near-memory cores. We differentiate the benefits stemming from each of elevating data access semantics, near-data computation, inter-accelerator coordination, and compute/access logic specialization. Experimental results indicate a geometric mean (energy efficiency improvement; speedup; data movement reduction) of (3.3; 1.59; 2.4), (2.46; 1.43; 3.5) and (1.46; 1.65; 1.48) compared to an out-of-order processor, monolithic accelerator with centralized accesses and monolithic accelerator with decentralized accesses, respectively. Evaluating both lightweight core and CGRA fabric implementations highlights model flexibility and quantifies the benefits of compute specialization for energy efficiency and speedup at 1.23 and 1.43, respectively.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- Infinity Stream: Portable and Programmer-Friendly In-/Near-Memory FusionZhengrong Wang, Christopher Liu, Aman Arora, Lizy Kurian John 等ASPLOS 2023 · 被引用 20 次
- Leviathan: A Unified System for General-Purpose Near-Data ComputingBrian C. Schwedock, Nathan BeckmannMICRO 2024 · 被引用 6 次
- AccelFlow: Orchestrating an On-Package Ensemble of Fine-Grained Accelerators for MicroservicesJovan Stojkovic, Abraham Farrell, Zhangxiaowen Gong, Christopher J. Hughes 等HPCA 2026 · 被引用 2 次
- NUPEA: Optimizing Critical Loads on Spatial Dataflow Architectures via Non-Uniform Processing-Element AccessSouradip Ghosh, Graham Gobieski, Keyi Zhang, Brandon Lucia 等ISCA 2025 · 被引用 1 次
相关 Paper
- Near-Stream Computing: General and Transparent Near-Cache AccelerationZhengrong Wang, Jian Weng, Sihao Liu, Tony NowatzkiHPCA 2022 · 被引用 24 次
- Compiler support for near data computingMahmut Taylan Kandemir, Jihyun Ryoo, Xulong Tang, Mustafa KaraköyPPoPP 2021 · 被引用 14 次
- AsyncDIMM: Achieving Asynchronous Execution in DIMM-Based Near-Memory ProcessingLiyan Chen, Dongxu Lyu, Jianfei Jiang, Qin Wang 等HPCA 2025 · 被引用 7 次
- täk¯: a polymorphic cache hierarchy for general-purpose optimization of data movementBrian C. Schwedock, Piratach Yoovidhya, Jennifer Seibert, Nathan BeckmannISCA 2022 · 被引用 16 次
- Livia: Data-Centric Computing Throughout the Memory HierarchyElliot Lockerman, Axel Feldmann, Mohammad Bakhshalipour, Alexandru Stanescu 等ASPLOS 2020 · 被引用 55 次
