Hyper-Ap: Enhancing Associative Processing Through A Full-Stack Optimization
Yue Zha, Jing Li
Abstract
Associative processing (AP) is a promising PIM paradigm that overcomes the von Neumann bottleneck (memory wall) by virtue of a radically different execution model. By decomposing arbitrary computations into a sequence of primitive memory operations (i.e., search and write), AP's execution model supports concurrent SIMD computations in-situ in the memory array to eliminate the need for data movement. This execution model also provides a native support for flexible data types and only requires a minimal modification on the existing memory design (low hardware complexity). Despite these advantages, the execution model of AP has two limitations that substantially increase the execution time, i.e., 1) it can only search a single pattern in one search operation and 2) it needs to perform a write operation after each search operation. In this paper, we propose the Highly Performant Associative Processor (Hyper- AP) to fully address the aforementioned limitations. The core of Hyper- AP is an enhanced execution model that reduces the number of search and write operations needed for computations, thereby reducing the execution time. This execution model is generic and improves the performance for both CMOS-based and RRAM-based AP, but it is more beneficial for the RRAMbased AP due to the substantially reduced write operations. We then provide complete architecture and micro-architecture with several optimizations to efficiently implement Hyper-AP. In order to reduce the programming complexity, we also develop a compilation framework so that users can write C-like programs with several constraints to run applications on Hyper- AP. Several optimizations have been applied in the compilation process to exploit the unique properties of Hyper- AP. Our experimental results show that, compared with the recent work IMP, Hyper- AP achieves up to 54×/4.4× better power-/area-efficiency for various representative arithmetic operations. For the evaluated benchmarks, Hyper-AP achieves 3.3× speedup and 23.8× energy reduction on average compared with IMP. Our evaluation also confirms that the proposed execution model is more beneficial for the RRAM-based AP than its CMOS-based counterpart.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 3667f1d9-1bca-4d4f-b6a1-980044a872f4Cited by top-tier papers7
- SIMDRAM: a framework for bit-serial SIMD processing using DRAMNastaran Hajinazar, Geraldo F. Oliveira, Sven Gregorio, João Dinis Ferreira et al.ASPLOS 2021 · 182 citations
- pLUTo: Enabling Massively Parallel Computation in DRAM via Lookup TablesJoão Dinis Ferreira, Gabriel Falcão, Juan Gómez-Luna, Mohammed Alser et al.MICRO 2022 · 60 citations
- CAPE: A Content-Addressable Processing EngineHelena Caminal, Kailin Yang, Srivatsa Srinivasa, Akshay Krishna Ramanathan et al.HPCA 2021 · 30 citations
- Accelerating database analytic query workloads using an associative processorHelena Caminal, Yannis Chronis, Tianshu Wu, Jignesh M. Patel et al.ISCA 2022 · 19 citations
- CATCAM: Constant-time Alteration Ternary CAM with Scalable In-Memory ArchitectureDibei Chen, Zhaoshi Li, Tianzhu Xiong, Zhiwei Liu et al.MICRO 2020 · 6 citations
Related papers
- FloatAP: Supporting High-Performance Floating-Point Arithmetic in Associative ProcessorsKailin Yang, José F. MartínezMICRO 2024 · 5 citations
- AmgR: Algebraic Multigrid Accelerated on ReRAMMingjia Fan, Xiaotian Tian, Yintao He, Junxian Li et al.DAC 2023 · 8 citations
- HH-PIM: Dynamic Optimization of Power and Performance with Heterogeneous-Hybrid PIM for Edge AI DevicesSangmin Jeon, Kangju Lee, Kyeongwon Lee, Woojoo LeeDAC 2025 · 4 citations
- BAAP: Coupling Compute-in-SRAM with DRAM Banks for Near-Memory ProcessingCecilio C. Tamarit, Socrates S. Wong, Akshati Vaishnav, José F. MartínezISCA 2026
- CINM (Cinnamon): A Compilation Infrastructure for Heterogeneous Compute In-Memory and Compute Near-Memory ParadigmsAsif Ali Khan, Hamid Farzaneh, Karl Friedrich Alexander Friebel, Clément Fournier et al.ASPLOS 2024 · 7 citations
