NUPEA: Optimizing Critical Loads on Spatial Dataflow Architectures via Non-Uniform Processing-Element Access
Souradip Ghosh, Graham Gobieski, Keyi Zhang, Brandon Lucia, Nathan Beckmann, Tony Nowatzki
摘要
Data movement is the dominant energy, performance, and scalability bottleneck in modern architectures. Systems have tackled data movement by distributing data, e.g., via non-uniform memory access (NUMA) architectures. However, to reduce data movement, these architectures must identify critical data and place it closer to compute. Clever data placement is complex and often ineffective.
Spatial dataflow architectures (SDAs) present a new opportunity to tackle data movement. SDAs distribute program instructions across a spatial fabric of processing elements (PEs). On large SDAs, some PEs are necessarily closer to memory than others, giving rise to non-uniform processing-element access (NUPEA). Clever instruction placement can thus reduce data movement by, e.g., placing critical loads close to memory.
This paper introduces NUPEA and contrasts it with prior datacentric approaches to scaling data movement. We find that it is often easier for the compiler to identify critical loads than the data they access, making NUPEA applicable where NUMA is not. We present simple architecture and compiler optimizations for NUPEA and implement them on the Monaco SDA architecture and effcc compiler, both industry products by Efficient Computer. On Monaco, across a range of important kernels, NUPEA yields an avg 28% speedup over a uniform-PE-access (UPEA) SDA and an avg 20% speed over a UPEA SDA with NUMA.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper24
- SIMDRAM: a framework for bit-serial SIMD processing using DRAMNastaran Hajinazar, Geraldo F. Oliveira, Sven Gregorio, João Dinis Ferreira 等ASPLOS 2021 · 被引用 182 次
- Snafu: An Ultra-Low-Power, Energy-Minimal CGRA-Generation Framework and ArchitectureGraham Gobieski, Ahmet Oguz Atli, Kenneth Mai, Brandon Lucia 等ISCA 2021 · 被引用 84 次
- A Hybrid Systolic-Dataflow Architecture for Inductive Matrix AlgorithmsJian Weng, Sihao Liu, Zhengrong Wang, Vidushi Dadu 等HPCA 2020 · 被引用 80 次
- Ultra-Elastic CGRAs for Irregular Loop SpecializationChristopher Torng, Peitian Pan, Yanghui Ou, Cheng Tan 等HPCA 2021 · 被引用 68 次
- A programmable, energy-minimal dataflow compiler and architectureGraham Gobieski, Souradip Ghosh, Marijn Heule, Todd C. Mowry 等MICRO 2022 · 被引用 66 次
相关 Paper
- TD-NUCA: Runtime Driven Management of NUCA Caches in Task Dataflow Programming ModelsPaul Caheny, Lluc Alvarez, Marc Casas, Miquel MoretóSC 2022 · 被引用 4 次
- Affinity Alloc: Taming Not-So Near-Data ComputingZhengrong Wang, Christopher Liu, Nathan Beckmann, Tony NowatzkiMICRO 2023 · 被引用 4 次
- MESA: Microarchitecture Extensions for Spatial Architecture GenerationDong Kai Wang, Jiaqi Lou, Naiyin Jin, Edwin Mascarenhas 等ISCA 2023 · 被引用 2 次
- Livia: Data-Centric Computing Throughout the Memory HierarchyElliot Lockerman, Axel Feldmann, Mohammad Bakhshalipour, Alexandru Stanescu 等ASPLOS 2020 · 被引用 55 次
- Near-Stream Computing: General and Transparent Near-Cache AccelerationZhengrong Wang, Jian Weng, Sihao Liu, Tony NowatzkiHPCA 2022 · 被引用 24 次
