TransPIM: A Memory-based Acceleration via Software-Hardware Co-Design for Transformer
Minxuan Zhou, Weihong Xu, Jaeyoung Kang, Tajana Rosing
摘要
Transformer-based models are state-of-the-art for many machine learning (ML) tasks. Executing Transformer usually requires a long execution time due to the large memory footprint and the low data reuse rate, stressing the memory system while under-utilizing the computing resources. Memory-based processing technologies, including processing in-memory (PIM) and near-memory computing (NMC), are promising to accelerate Transformer since they provide high memory bandwidth utilization and extensive computation parallelism. However, the previous memory-based ML accelerators mainly target at optimizing dataflow and hardware for compute-intensive ML models (e.g., CNNs), which do not fit the memory-intensive characteristics of Transformer. In this work, we propose TransPIM, a memory-based acceleration for Transformer using software and hardware co-design. In the software-level, TransPIM adopts a token-based dataflow to avoid the expensive inter-layer data movements introduced by previous layer-based dataflow. In the hardware-level, TransPIM introduces lightweight modifications in the conventional high bandwidth memory (HBM) architecture to support PIM-NMC hybrid processing and efficient data communication for accelerating Transformer-based models. Our experiments show that TransPIM is 3.7× to 9.1× faster than existing memory-based acceleration. As compared to conventional accelerators, TransPIM is 22.1× to 114.9× faster than GPUs and provides 2.0× more throughput than existing ASIC-based accelerators.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper23
- NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM InferencingGuseul Heo, Sangyeop Lee, Jaehong Cho, Hyunmin Choi 等ASPLOS 2024 · 被引用 121 次
- IANUS: Integrated Accelerator based on NPU-PIM Unified Memory SystemMinseok Seo, Xuan Truong Nguyen, Seok Joong Hwang, Yongkee Kwon 等ASPLOS 2024 · 被引用 57 次
- PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model InferenceYufeng Gu, Alireza Khadem, Sumanth Umesh, Ning Liang 等ASPLOS 2025 · 被引用 44 次
- Lightening-Transformer: A Dynamically-Operated Optically-Interconnected Photonic Transformer AcceleratorHanqing Zhu, Jiaqi Gu, Hanrui Wang, Zixuan Jiang 等HPCA 2024 · 被引用 44 次
- Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous BatchingSungmin Yun, Kwanhee Kyung, Juhwan Cho, Jaewan Choi 等MICRO 2024 · 被引用 40 次
相关 Paper
- HAIMA: A Hybrid SRAM and DRAM Accelerator-in-Memory Architecture for TransformerYan Ding, Chubo Liu, Mingxing Duan, Wanli Chang 等DAC 2023 · 被引用 24 次
- AttAcc! Unleashing the Power of PIM for Batched Transformer-based Generative Model InferenceJaehyun Park, Jaewan Choi, Kwanhee Kyung, Michael Jaemin Kim 等ASPLOS 2024 · 被引用 125 次
- A Real-time Execution System of Multimodal Transformer through PIM-GPU CollaborationShengyi Ji, Chubo Liu, Yan Ding, Qing Liao 等DAC 2024 · 被引用 2 次
- A Convolution Neural Network Accelerator Design with Weight Mapping and Pipeline OptimizationLixia Han, Peng Huang, Zheng Zhou, Yiyang Chen 等DAC 2023 · 被引用 8 次
- PAISE: PIM-Accelerated Inference Scheduling Engine for Transformer-based LLMHyojung Lee, Daehyeon Baek, Jimyoung Son, Jieun Choi 等HPCA 2025 · 被引用 10 次
