Pimba: A Processing-in-Memory Acceleration for Post-Transformer Large Language Model Serving
Wonung Kim, Yubin Lee, Yoonsung Kim, Jinwoo Hwang, Seongryong Oh, Jiyong Jung, Aziz Huseynov, Woong Gyu Park, Chang Hyun Park, Divya Mahajan, Jongse Park
摘要
Transformers are the driving force behind today's Large Language Models (LLMs), serving as the foundation for their performance and versatility. Yet, their compute and memory costs grow with sequence length, posing scalability challenges for long-context inferencing. In response, the algorithm community is exploring alternative architectures-such as state space models (SSMs) (e.g., Mamba-2), linear attention, and recurrent neural networks (RNNs)-which we refer to as post-transformers. This shift presents a key challenge: building a serving system that efficiently supports not only emerging post-transformer LLMs but also existing transformer models within a unified framework.
To address this challenge, we analyze the performance characteristics of transformer and post-transformer LLMs. Despite their algorithmic differences, both are largely bounded by memory bandwidth under batched inference-due to attention in transformers and state updates in post-transformers. Inspired by this finding, we propose Pimba, an accelerator solution that aims to address the memory bottleneck by jointly leveraging (1) Processing-in-Memory (PIM) paradigm and (2) LLM quantization. Further analyses suggest two additional insights: (1) state update operations, unlike attention, incur high hardware cost, making per-bank PIM acceleration inefficient, and (2) different low-precision arithmetic methods offer varying accuracy-area tradeoffs, while we identify Microsoft's MX as a Pareto-optimal choice. Building on these insights, we design the
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- -LLM: An Integrated NPU-PIM Accelerator for Edge LLM Inference Using Hybrid Numerical FormatsYuzong Chen, Chao Fang, Xilai Dai, Yuheng Wu 等ISCA 2026 · 被引用 4 次
- DCC: Data-Centric Compilation of Machine Learning Kernels for Processing-In-Memory ArchitecturesPeiming Yang, Sankeerth Durvasula, Ivan Fernandez, Mohammad Sadrosadati 等ISCA 2026 · 被引用 3 次
它引用的顶会 Paper38
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
相关 Paper
- Mamba-3: Improved Sequence Modeling using State Space PrinciplesAakash Sunil Lahoti, Kevin Y. Li, Berlin Chen, Caitlin Wang 等ICLR 2026 · 被引用 96 次
- MambaExtend: A Training-Free Approach to Improve Long Context Extension of MambaSeyedarmin Azizi, Souvik Kundu, Mohammad Erfan Sadeghi, Massoud PedramICLR 2025
- Quamba: A Post-Training Quantization Recipe for Selective State Space ModelsHung-Yueh Chiang, Chi-Chih Chang, Natalia Frumkin, Kai-Chiang Wu 等ICLR 2025
- The Expressive Capacity of State Space Models: A Formal Language PerspectiveYash Raj Sarrof, Yana Veitsman, Michael HahnNeurIPS 2024 · 被引用 53 次
- HLX: A Unified Pipelined Architecture for Optimized Performance of Hybrid Transformer-Mamba Language ModelsIn-Jun Jung, Gyeongrok Yang, Jaeha Min, Joo-Young KimMICRO 2025 · 被引用 1 次
