A Real-time Execution System of Multimodal Transformer through PIM-GPU Collaboration
Shengyi Ji, Chubo Liu, Yan Ding, Qing Liao, Zhuo Tang
Abstract
Multimodal transformer excels in various applications, but faces great challenges such as high memory consumption and limited data reuse that hinder real-time performance. To address these issues, we propose a processing-in-memory (PIM)-GPU collaboration oriented compiler to accelerate the multimodal transformers. The PIM-GPU collaboration adapts well to multimodal transformers and significantly accelerates model inference. In addition, we introduce a tailored PIM allocation algorithm for variable-length inputs to further improve computation efficiency. Experimental results show that our scheme can achieve an average 15x end-to-end speedup.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 82ae2940-9ec4-487e-8876-74a757cf3136Related papers
- TransPIM: A Memory-based Acceleration via Software-Hardware Co-Design for TransformerMinxuan Zhou, Weihong Xu, Jaeyoung Kang, Tajana RosingHPCA 2022 · 142 citations
- IANUS: Integrated Accelerator based on NPU-PIM Unified Memory SystemMinseok Seo, Xuan Truong Nguyen, Seok Joong Hwang, Yongkee Kwon et al.ASPLOS 2024 · 57 citations
- DCC: Data-Centric Compilation of Machine Learning Kernels for Processing-In-Memory ArchitecturesPeiming Yang, Sankeerth Durvasula, Ivan Fernandez, Mohammad Sadrosadati et al.ISCA 2026 · 3 citations
- PIMoE: Towards Efficient MoE Transformer Deployment on NPU-PIM System through Throttle-Aware Task OffloadingLizhou Wu, Haozhe Zhu, Siqi He, Xuanda Lin et al.DAC 2025 · 6 citations
- PIMphony: Overcoming Bandwidth and Capacity Inefficiency in PIM-Based Long-Context LLM Inference SystemHyucksung Kwon, Kyungmo Koo, Janghyeon Kim, Woongkyu Lee et al.HPCA 2026 · 3 citations
