NUPEA: Optimizing Critical Loads on Spatial Dataflow Architectures via Non-Uniform Processing-Element Access
Souradip Ghosh, Graham Gobieski, Keyi Zhang, Brandon Lucia, Nathan Beckmann, Tony Nowatzki
Abstract
Data movement is the dominant energy, performance, and scalability bottleneck in modern architectures. Systems have tackled data movement by distributing data, e.g., via non-uniform memory access (NUMA) architectures. However, to reduce data movement, these architectures must identify critical data and place it closer to compute. Clever data placement is complex and often ineffective.
Spatial dataflow architectures (SDAs) present a new opportunity to tackle data movement. SDAs distribute program instructions across a spatial fabric of processing elements (PEs). On large SDAs, some PEs are necessarily closer to memory than others, giving rise to non-uniform processing-element access (NUPEA). Clever instruction placement can thus reduce data movement by, e.g., placing critical loads close to memory.
This paper introduces NUPEA and contrasts it with prior datacentric approaches to scaling data movement. We find that it is often easier for the compiler to identify critical loads than the data they access, making NUPEA applicable where NUMA is not. We present simple architecture and compiler optimizations for NUPEA and implement them on the Monaco SDA architecture and effcc compiler, both industry products by Efficient Computer. On Monaco, across a range of important kernels, NUPEA yields an avg 28% speedup over a uniform-PE-access (UPEA) SDA and an avg 20% speed over a UPEA SDA with NUMA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 73e2b7da-79f6-4c1a-836a-69a01fbc588eCited by top-tier papers1
Ask how each one uses itBuilds on24
- SIMDRAM: a framework for bit-serial SIMD processing using DRAMNastaran Hajinazar, Geraldo F. Oliveira, Sven Gregorio, João Dinis Ferreira et al.ASPLOS 2021 · 182 citations
- Snafu: An Ultra-Low-Power, Energy-Minimal CGRA-Generation Framework and ArchitectureGraham Gobieski, Ahmet Oguz Atli, Kenneth Mai, Brandon Lucia et al.ISCA 2021 · 84 citations
- A Hybrid Systolic-Dataflow Architecture for Inductive Matrix AlgorithmsJian Weng, Sihao Liu, Zhengrong Wang, Vidushi Dadu et al.HPCA 2020 · 80 citations
- Ultra-Elastic CGRAs for Irregular Loop SpecializationChristopher Torng, Peitian Pan, Yanghui Ou, Cheng Tan et al.HPCA 2021 · 68 citations
- A programmable, energy-minimal dataflow compiler and architectureGraham Gobieski, Souradip Ghosh, Marijn Heule, Todd C. Mowry et al.MICRO 2022 · 66 citations
Related papers
- TD-NUCA: Runtime Driven Management of NUCA Caches in Task Dataflow Programming ModelsPaul Caheny, Lluc Alvarez, Marc Casas, Miquel MoretóSC 2022 · 4 citations
- Affinity Alloc: Taming Not-So Near-Data ComputingZhengrong Wang, Christopher Liu, Nathan Beckmann, Tony NowatzkiMICRO 2023 · 4 citations
- MESA: Microarchitecture Extensions for Spatial Architecture GenerationDong Kai Wang, Jiaqi Lou, Naiyin Jin, Edwin Mascarenhas et al.ISCA 2023 · 2 citations
- Livia: Data-Centric Computing Throughout the Memory HierarchyElliot Lockerman, Axel Feldmann, Mohammad Bakhshalipour, Alexandru Stanescu et al.ASPLOS 2020 · 55 citations
- Near-Stream Computing: General and Transparent Near-Cache AccelerationZhengrong Wang, Jian Weng, Sihao Liu, Tony NowatzkiHPCA 2022 · 24 citations
