OptiPIM: Optimizing Processing-in-Memory Acceleration Using Integer Linear Programming
Jiantao Liu, Minxuan Zhou, Yue Pan, Chien-Yi Yang, Lana Josipovic, Tajana Rosing
Abstract
Processing-in-memory (PIM) accelerators provide superior performance and energy efficiency to conventional architectures by minimizing off-chip data movement and exploiting extensive internal memory bandwidth for computation.However, efficient PIM acceleration requires careful software-hardware mapping that transforms application algorithms into PIM operations and data layout.Unfortunately, existing PIM accelerators adopt manually tuned heuristics or exhaustive search to determine the mappings on PIM accelerators, leading to under-optimized performance and/or long optimization time.In this work, we propose OptiPIM, a novel optimization framework based on Integer Linear Programming (ILP) to efficiently generate the optimal mapping for data-intensive applications on PIM accelerators.The proposed framework adopts a PIM-friendly mapping representation with accurate cost modeling and a concise description of the entire design space, allowing us to formulate an efficient and effective ILP problem and optimize the mapping on PIM architectures.We implement OptiPIM in the opensource MLIR framework, enabling OptiPIM to generate optimized mappings for PyTorch workloads on PIM accelerators.We evaluate widely used machine learning workloads on two state-of-the-art PIM accelerators.Our experiments show that OptiPIM can generate optimal mappings within 4 minutes.Mappings generated by OptiPIM are at least 1.9× faster than those generated by heuristics.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 50bfec9c-92e3-4de6-824f-655379ccc92bCited by top-tier papers4
- PuDghost: Experimental Analysis of Computation Result Corruption in Processing-Using-Dram Operations on Real Dram Chips and Implications for Future SystemsDaichi Tokuda, Ismail Emir Yüksel, Tatsuya Kubo, Ataberk Olgun et al.ISCA 2026 · 4 citations
- DCC: Data-Centric Compilation of Machine Learning Kernels for Processing-In-Memory ArchitecturesPeiming Yang, Sankeerth Durvasula, Ivan Fernandez, Mohammad Sadrosadati et al.ISCA 2026 · 3 citations
- Hardwired-Neuron Language Processing Units as General-Purpose Cognitive SubstratesYang Liu, Yi Chen, Yongwei Zhao, Yifan Hao et al.ASPLOS 2026
- Bringing Near Data Processing Into the Low-Bit Floating-Point EraTongxin Xie, Mingyu Gao, Zehao Wang, Zhihao Jia et al.ISCA 2026
Related papers
- Chimera: An Analytical Optimizing Framework for Effective Compute-intensive Operators FusionSize Zheng, Siyuan Chen, Peidi Song, Renze Chen et al.HPCA 2023 · 46 citations
- PIM-MMU: A Memory Management Unit for Accelerating Data Transfers in Commercial PIM SystemsDongjae Lee, Bongjoon Hyun, Taehun Kim, Minsoo RhuMICRO 2024 · 23 citations
- To PIM or not for emerging general purpose processing in DDR memory systemsAlexandar Devic, Siddhartha Balakrishna Rai, Anand Sivasubramaniam, Ameen Akel et al.ISCA 2022 · 51 citations
- ML-CGRA: An Integrated Compilation Framework to Enable Efficient Machine Learning Acceleration on CGRAsYixuan Luo, Cheng Tan, Nicolas Bohm Agostini, Ang Li et al.DAC 2023 · 39 citations
- An Optimizing Framework on MLIR for Efficient FPGA-based Accelerator GenerationWeichuang Zhang, Jieru Zhao, Guan Shen, Quan Chen et al.HPCA 2024 · 8 citations
