Infinity Stream: Portable and Programmer-Friendly In-/Near-Memory Fusion
Zhengrong Wang, Christopher Liu, Aman Arora, Lizy Kurian John, Tony Nowatzki
Abstract
In-memory computing with large last-level caches is promising to dramatically alleviate data movement bottlenecks and expose massive bitline-level parallelization opportunities. However, key challenges from its unique execution model remain unsolved: automated parallelization, transparently orchestrating data transposition/alignment/broadcast for bit-serial logic, and mixing in-/nearmemory computing. Most importantly, the solution should be programmer friendly and portable across platforms.
Our key innovation is an execution model and intermediate representation (IR) that enables hybrid CPU-core, in-memory, and nearmemory processing. Our IR is the tensor dataflow graph (tDFG), which is a unified representation of in-memory and near-memory computation. The tDFG exposes tensor-data structure information so that the hardware and runtime can automatically orchestrate data management for bit-serial execution, including runtime data layout transformations. To enable microarchitecture portability, we use a two-phase, JIT-based compilation approach to dynamically lower the tDFG to in-memory commands.
Our design, infinity stream, is evaluated on a cycle-accurate simulator. Across data-processing workloads with fp32, it achieves 2.6× speedup and 75% traffic reduction over a state-of-the-art nearmemory computing technique, with 2.4× energy efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 86f5acb5-0307-42ab-9ae1-150f911bfc65Cited by top-tier papers9
- NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM InferencingGuseul Heo, Sangyeop Lee, Jaehong Cho, Hyunmin Choi et al.ASPLOS 2024 · 121 citations
- BP-NTT: Fast and Compact in-SRAM Number Theoretic Transform with Bit-Parallel Modular MultiplicationJingyao Zhang, Mohsen Imani, Elaheh SadrediniDAC 2023 · 23 citations
- CINM (Cinnamon): A Compilation Infrastructure for Heterogeneous Compute In-Memory and Compute Near-Memory ParadigmsAsif Ali Khan, Hamid Farzaneh, Karl Friedrich Alexander Friebel, Clément Fournier et al.ASPLOS 2024 · 7 citations
- Leviathan: A Unified System for General-Purpose Near-Data ComputingBrian C. Schwedock, Nathan BeckmannMICRO 2024 · 6 citations
- Stream-Based Data Placement for Near-Data Processing with Extended MemoryYiwei Li, Boyu Tian, Yi Ren, Mingyu GaoMICRO 2024 · 5 citations
Builds on21
- SIMDRAM: a framework for bit-serial SIMD processing using DRAMNastaran Hajinazar, Geraldo F. Oliveira, Sven Gregorio, João Dinis Ferreira et al.ASPLOS 2021 · 182 citations
- egg: Fast and extensible equality saturationMax Willsey, Chandrakana Nandi, Yisu Remy Wang, Oliver Flatt et al.POPL 2021 · 170 citations
- DSAGEN: Synthesizing Programmable Spatial AcceleratorsJian Weng, Sihao Liu, Vidushi Dadu, Zhengrong Wang et al.ISCA 2020 · 140 citations
- SpaceA: Sparse Matrix Vector Multiplication on Processing-in-Memory AcceleratorXinfeng Xie, Zheng Liang, Peng Gu, Abanti Basak et al.HPCA 2021 · 111 citations
- DUAL: Acceleration of Clustering Algorithms using Digital-based Processing In-MemoryMohsen Imani, Saikishan Pampana, Saransh Gupta, Minxuan Zhou et al.MICRO 2020 · 91 citations
Related papers
- Near-Stream Computing: General and Transparent Near-Cache AccelerationZhengrong Wang, Jian Weng, Sihao Liu, Tony NowatzkiHPCA 2022 · 24 citations
- StreamTensor: Make Tensors Stream in Dataflow Accelerators for LLMsHanchen Ye, Deming ChenMICRO 2025 · 5 citations
- TD-NUCA: Runtime Driven Management of NUCA Caches in Task Dataflow Programming ModelsPaul Caheny, Lluc Alvarez, Marc Casas, Miquel MoretóSC 2022 · 4 citations
- VTC: DNN Compilation with Virtual Tensors for Data Movement EliminationMuyan Hu, Ahan Gupta, Jiachen Yuan, Vima Gupta et al.OSDI 2026
- CONTINUUM: Restoring the Contiguous Tensor Abstraction Efficiently for Dynamic AI Workloads via Hardware VirtualizationYangyu Zhang, shuoming zhang, Chunwei Xia, Shuaijiang Li et al.ICML 2026
