Near-Stream Computing: General and Transparent Near-Cache Acceleration
Zhengrong Wang, Jian Weng, Sihao Liu, Tony Nowatzki
Abstract
Data movement and communication have become the primary bottlenecks in large multicore systems. The near-data computing paradigm provides a solution: move computation to where the data resides on-chip. Two challenges keep near-data computing from the mainstream: lack of programmer transparency and applicability. Programmer transparency requires providing sequential memory semantics with distributed computation, which requires burdensome coordination. Broad applicability requires support for combinations of address patterns (e.g. affine, indirect, multi-operand) and computation types (loads, stores, reductions, atomics).
We find that streams -coarse grain memory access patterns -are a powerful ISA abstraction for near data offloading.
Tracking data access at stream-granularity heavily reduces the burden of coordination for providing sequential semantics. Decomposing the problem using streams means that arbitrary combinations of address and computation patterns can be combined for broad generality.
With this insight, we develop a paradigm called near-stream computing, comprising a compiler, CPU ISA extension, and a microarchitecture that facilitate programmer transparent computation offloading to shared caches. We evaluate our system on OpenMP kernels that stress broad addressing and compute behavior, and find that 46% of dynamic instructions can be offloaded to remote banks, reducing the network traffic by 76%. Overall it achieves 2.13× speedup over a state-ofthe-art near-data computing technique, with a 1.90× energy efficiency gain.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ccf00324-9d45-4043-a0e9-7f0d58dab5f5Cited by top-tier papers9
- OverGen: Improving FPGA Usability through Domain-specific Overlay GenerationSihao Liu, Jian Weng, Dylan Kupsh, Atefeh Sohrabizadeh et al.MICRO 2022 · 32 citations
- TaskStream: accelerating task-parallel workloads by recovering program structureVidushi Dadu, Tony NowatzkiASPLOS 2022 · 23 citations
- Infinity Stream: Portable and Programmer-Friendly In-/Near-Memory FusionZhengrong Wang, Christopher Liu, Aman Arora, Lizy Kurian John et al.ASPLOS 2023 · 20 citations
- Leviathan: A Unified System for General-Purpose Near-Data ComputingBrian C. Schwedock, Nathan BeckmannMICRO 2024 · 6 citations
- Stream-Based Data Placement for Near-Data Processing with Extended MemoryYiwei Li, Boyu Tian, Yi Ren, Mingyu GaoMICRO 2024 · 5 citations
Builds on16
- DSAGEN: Synthesizing Programmable Spatial AcceleratorsJian Weng, Sihao Liu, Vidushi Dadu, Zhengrong Wang et al.ISCA 2020 · 140 citations
- SpaceA: Sparse Matrix Vector Multiplication on Processing-in-Memory AcceleratorXinfeng Xie, Zheng Liang, Peng Gu, Abanti Basak et al.HPCA 2021 · 111 citations
- DUAL: Acceleration of Clustering Algorithms using Digital-based Processing In-MemoryMohsen Imani, Saikishan Pampana, Saransh Gupta, Minxuan Zhou et al.MICRO 2020 · 91 citations
- FAFNIR: Accelerating Sparse Gathering by Using Efficient Near-Memory Intelligent ReductionBahar Asgari, Ramyad Hadidi, Jiashen Cao, Da Eun Shim et al.HPCA 2021 · 87 citations
- A Hybrid Systolic-Dataflow Architecture for Inductive Matrix AlgorithmsJian Weng, Sihao Liu, Zhengrong Wang, Vidushi Dadu et al.HPCA 2020 · 80 citations
Related papers
- Affinity Alloc: Taming Not-So Near-Data ComputingZhengrong Wang, Christopher Liu, Nathan Beckmann, Tony NowatzkiMICRO 2023 · 4 citations
- An architecture interface and offload model for low-overhead, near-data, distributed acceleratorsSaambhavi Baskaran, Mahmut Taylan Kandemir, Jack SampsonMICRO 2022 · 13 citations
- Stream Floating: Enabling Proactive and Decentralized Cache OptimizationsZhengrong Wang, Jian Weng, Jason Lowe-Power, Jayesh Gaur et al.HPCA 2021 · 27 citations
- Conduit: Programmer-Transparent Near-Data Processing Using Multiple Compute-Capable Resources in Solid State DrivesRakesh Nadig, Vamanan Arulchelvan, Mayank Kabra, Harshita Gupta et al.HPCA 2026 · 2 citations
- Compiler support for near data computingMahmut Taylan Kandemir, Jihyun Ryoo, Xulong Tang, Mustafa KaraköyPPoPP 2021 · 14 citations
