Near-Stream Computing: General and Transparent Near-Cache Acceleration
Zhengrong Wang, Jian Weng, Sihao Liu, Tony Nowatzki
摘要
Data movement and communication have become the primary bottlenecks in large multicore systems. The near-data computing paradigm provides a solution: move computation to where the data resides on-chip. Two challenges keep near-data computing from the mainstream: lack of programmer transparency and applicability. Programmer transparency requires providing sequential memory semantics with distributed computation, which requires burdensome coordination. Broad applicability requires support for combinations of address patterns (e.g. affine, indirect, multi-operand) and computation types (loads, stores, reductions, atomics).
We find that streams -coarse grain memory access patterns -are a powerful ISA abstraction for near data offloading.
Tracking data access at stream-granularity heavily reduces the burden of coordination for providing sequential semantics. Decomposing the problem using streams means that arbitrary combinations of address and computation patterns can be combined for broad generality.
With this insight, we develop a paradigm called near-stream computing, comprising a compiler, CPU ISA extension, and a microarchitecture that facilitate programmer transparent computation offloading to shared caches. We evaluate our system on OpenMP kernels that stress broad addressing and compute behavior, and find that 46% of dynamic instructions can be offloaded to remote banks, reducing the network traffic by 76%. Overall it achieves 2.13× speedup over a state-ofthe-art near-data computing technique, with a 1.90× energy efficiency gain.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- OverGen: Improving FPGA Usability through Domain-specific Overlay GenerationSihao Liu, Jian Weng, Dylan Kupsh, Atefeh Sohrabizadeh 等MICRO 2022 · 被引用 32 次
- TaskStream: accelerating task-parallel workloads by recovering program structureVidushi Dadu, Tony NowatzkiASPLOS 2022 · 被引用 23 次
- Infinity Stream: Portable and Programmer-Friendly In-/Near-Memory FusionZhengrong Wang, Christopher Liu, Aman Arora, Lizy Kurian John 等ASPLOS 2023 · 被引用 20 次
- Leviathan: A Unified System for General-Purpose Near-Data ComputingBrian C. Schwedock, Nathan BeckmannMICRO 2024 · 被引用 6 次
- Stream-Based Data Placement for Near-Data Processing with Extended MemoryYiwei Li, Boyu Tian, Yi Ren, Mingyu GaoMICRO 2024 · 被引用 5 次
它引用的顶会 Paper16
- DSAGEN: Synthesizing Programmable Spatial AcceleratorsJian Weng, Sihao Liu, Vidushi Dadu, Zhengrong Wang 等ISCA 2020 · 被引用 140 次
- SpaceA: Sparse Matrix Vector Multiplication on Processing-in-Memory AcceleratorXinfeng Xie, Zheng Liang, Peng Gu, Abanti Basak 等HPCA 2021 · 被引用 111 次
- DUAL: Acceleration of Clustering Algorithms using Digital-based Processing In-MemoryMohsen Imani, Saikishan Pampana, Saransh Gupta, Minxuan Zhou 等MICRO 2020 · 被引用 91 次
- FAFNIR: Accelerating Sparse Gathering by Using Efficient Near-Memory Intelligent ReductionBahar Asgari, Ramyad Hadidi, Jiashen Cao, Da Eun Shim 等HPCA 2021 · 被引用 87 次
- A Hybrid Systolic-Dataflow Architecture for Inductive Matrix AlgorithmsJian Weng, Sihao Liu, Zhengrong Wang, Vidushi Dadu 等HPCA 2020 · 被引用 80 次
相关 Paper
- Affinity Alloc: Taming Not-So Near-Data ComputingZhengrong Wang, Christopher Liu, Nathan Beckmann, Tony NowatzkiMICRO 2023 · 被引用 4 次
- An architecture interface and offload model for low-overhead, near-data, distributed acceleratorsSaambhavi Baskaran, Mahmut Taylan Kandemir, Jack SampsonMICRO 2022 · 被引用 13 次
- Stream Floating: Enabling Proactive and Decentralized Cache OptimizationsZhengrong Wang, Jian Weng, Jason Lowe-Power, Jayesh Gaur 等HPCA 2021 · 被引用 27 次
- Conduit: Programmer-Transparent Near-Data Processing Using Multiple Compute-Capable Resources in Solid State DrivesRakesh Nadig, Vamanan Arulchelvan, Mayank Kabra, Harshita Gupta 等HPCA 2026 · 被引用 2 次
- Compiler support for near data computingMahmut Taylan Kandemir, Jihyun Ryoo, Xulong Tang, Mustafa KaraköyPPoPP 2021 · 被引用 14 次
