IANUS: Integrated Accelerator based on NPU-PIM Unified Memory System
Minseok Seo, Xuan Truong Nguyen, Seok Joong Hwang, Yongkee Kwon, Guhyun Kim, Chanwook Park, Ilkon Kim, Jaehan Park, Jeongbin Kim, Woojae Shin, Jongsoon Won, Haerang Choi
摘要
Accelerating end-to-end inference of transformer-based large language models (LLMs) is a critical component of AI services in datacenters. However, the diverse compute characteristics of LLMs' end-to-end inference present challenges as previously proposed accelerators only address certain operations or stages (e.g., self-attention, generation stage, etc.). To address the unique challenges of accelerating end-to-end inference, we propose IANUS - Integrated Accelerator based on NPU-PIM Unified Memory System. IANUS is a domain-specific system architecture that combines a Neural Processing Unit (NPU) with a Processing-in-Memory (PIM) to leverage both the NPU's high computation throughput and the PIM's high effective memory bandwidth. In particular, IANUS employs a unified main memory system where the PIM memory is used both for PIM operations and for NPU's main memory. The unified main memory system ensures that memory capacity is efficiently utilized and the movement of shared data between NPU and PIM is minimized. However, it introduces new challenges since normal memory accesses and PIM computations cannot be performed simultaneously. Thus, we propose novel PIM Access Scheduling that manages not only the scheduling of normal memory accesses and PIM computations but also workload mapping across the PIM and the NPU. Our detailed simulation evaluations show that IANUS improves the performance of GPT-2 by 6.2× and 3.2×, on average, compared to the NVIDIA A100 GPU and the state-of-the-art accelerator. As a proof-of-concept, we develop a prototype of IANUS with a commercial PIM, NPU, and an FPGA-based PIM controller to demonstrate the feasibility of IANUS.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- PAPI: Exploiting Dynamic Parallelism in Large Language Model Decoding with a Processing-In-Memory-Enabled Computing SystemYintao He, Haiyu Mao, Christina Giannoula, Mohammad Sadrosadati 等ASPLOS 2025 · 被引用 37 次
- Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache QuantizationMinsu Kim, Seongmin Hong, Ryeowook Ko, Soongyu Choi 等ISCA 2025 · 被引用 17 次
- Make LLM Inference Affordable to Everyone: Augmenting GPU Memory with NDP-DIMMLian Liu, Shixin Zhao, Bing Li, Haimeng Ren 等HPCA 2025 · 被引用 15 次
- UniNDP: A Unified Compilation and Simulation Tool for Near DRAM Processing ArchitecturesTongxin Xie, Zhenhua Zhu, Bing Li, Yukai He 等HPCA 2025 · 被引用 9 次
- PIMnet: A Domain-Specific Network for Efficient Collective Communication in Scalable PIMHyojun Son, Gilbert Jonatan, Xiangyu Wu, Haeyoon Cho 等HPCA 2025 · 被引用 7 次
它引用的顶会 Paper9
- SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head PruningHanrui Wang, Zhekai Zhang, Song HanHPCA 2021 · 被引用 412 次
- Sanger: A Co-Design Framework for Enabling Sparse Attention using Reconfigurable ArchitectureLiqiang Lu, Yicheng Jin, Hangrui Bi, Zizhang Luo 等MICRO 2021 · 被引用 221 次
- ELSA: Hardware-Software Co-design for Efficient, Lightweight Self-Attention Mechanism in Neural NetworksTae Jun Ham, Yejin Lee, Seong Hoon Seo, Soosung Kim 等ISCA 2021 · 被引用 185 次
- TransPIM: A Memory-based Acceleration via Software-Hardware Co-Design for TransformerMinxuan Zhou, Weihong Xu, Jaeyoung Kang, Tajana RosingHPCA 2022 · 被引用 142 次
- TurboTransformers: an efficient GPU serving system for transformer modelsJiarui Fang, Yang Yu, Chengduo Zhao, Jie ZhouPPoPP 2021 · 被引用 117 次
相关 Paper
- NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM InferencingGuseul Heo, Sangyeop Lee, Jaehong Cho, Hyunmin Choi 等ASPLOS 2024 · 被引用 121 次
- AttAcc! Unleashing the Power of PIM for Batched Transformer-based Generative Model InferenceJaehyun Park, Jaewan Choi, Kwanhee Kyung, Michael Jaemin Kim 等ASPLOS 2024 · 被引用 125 次
- PAISE: PIM-Accelerated Inference Scheduling Engine for Transformer-based LLMHyojung Lee, Daehyeon Baek, Jimyoung Son, Jieun Choi 等HPCA 2025 · 被引用 10 次
- PIMphony: Overcoming Bandwidth and Capacity Inefficiency in PIM-Based Long-Context LLM Inference SystemHyucksung Kwon, Kyungmo Koo, Janghyeon Kim, Woongkyu Lee 等HPCA 2026 · 被引用 3 次
- BlockPIM: Optimizing Memory Management for PIM-enabled Long-Context LLM InferenceZhichun Li, Jun Zhou, Xueqi Li, Ninghui SunDAC 2025 · 被引用 3 次
