Sparse Attention Acceleration with Synergistic In-Memory Pruning and On-Chip Recomputation
Amir Yazdanbakhsh, Ashkan Moradifirouzabadi, Zheng Li, Mingu Kang
摘要
comparisons against threshold values can outweigh the benefits of in-memory computing.
- Selective read of unpruned embeddings: Supporting inmemory ReRAM pruning enforces a particular data layout for key embeddings. However, this layout constraints the ability to selectively read the unpruned vectors. To remedy these considerations, this work makes the following contributions: 1 We introduce a unique perspective on the ReRAM in-memory computing paradigm. We employ approximate in-memory compute and precise on-chip recompute in tandem to mitigate the likely negative repercussions to model accuracy due to inherent circuit inaccuracies. 2 We employ analog comparators to carry out the comparisons with threshold values and instead produce 1-bit data to indicate the pruning status. With this shift in design, we reduce the hardware cost, which is proportional to input bit precision, to merely the cost of a series of 1-bit analog to digital converters (ADCs). 3 We repurpose an existing solution, which enables us to implement data reuse based on our observations. On the hardware side, we rely on recently taped-out transposable ReRAMs [141] that introduce in-situ transposed read access. While initially intended for efficiently accessing neural network weights, our application of this hardware selectively reads unpruned embeddings. For the data reuse, we observe that there is a considerable spatial locality between unpruned key vectors of adjacent queries. We exploit this spatial location to improve data reuse and further reduce the data communication overhead.
We evaluate our approach in several self-attention models with large sequences, including BERT, ALBERT, ViT, GPT-2, and two futuristic designs (e.g. 2K and 4K input sequence length). Under an iso design, our results show that, on average, SPRINT delivers 7.5× speed-up and 19.6× energy reduction compared to a baseline design with 16KB on-chip memory. The benefit increases as on-chip resources become scarcer, representing a design point for resource constrained platforms, e.g. 1.6× more energy reduction with 16KB on-chip memory than the case with 64KB capacity.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- FLAT: An Optimized Dataflow for Mitigating Attention BottlenecksSheng-Chun Kao, Suvinay Subramanian, Gaurav Agrawal, Amir Yazdanbakhsh 等ASPLOS 2023 · 被引用 68 次
- IANUS: Integrated Accelerator based on NPU-PIM Unified Memory SystemMinseok Seo, Xuan Truong Nguyen, Seok Joong Hwang, Yongkee Kwon 等ASPLOS 2024 · 被引用 57 次
- RAELLA: Reforming the Arithmetic for Efficient, Low-Resolution, and Low-Loss Analog PIM: No Retraining Required!Tanner Andrulis, Joel S. Emer, Vivienne SzeISCA 2023 · 被引用 45 次
- PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model InferenceYufeng Gu, Alireza Khadem, Sumanth Umesh, Ning Liang 等ASPLOS 2025 · 被引用 44 次
- Tender: Accelerating Large Language Models via Tensor Decomposition and Runtime RequantizationJungi Lee, Wonbeom Lee, Jaewoong SimISCA 2024 · 被引用 41 次
它引用的顶会 Paper26
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- Long Range Arena : A Benchmark for Efficient TransformersYi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen 等ICLR 2021 · 被引用 881 次
- Compressive Transformers for Long-Range Sequence ModellingJack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier 等ICLR 2020 · 被引用 833 次
相关 Paper
- Look-Up Table based Energy Efficient Processing in Cache Support for Neural Network AccelerationAkshay Krishna Ramanathan, Gurpreet S. Kalsi, Srivatsa Srinivasa, Tarun Makesh Chandran 等MICRO 2020 · 被引用 50 次
- Timely: Pushing Data Movements And Interfaces In Pim Accelerators Towards Local And In Time DomainWeitao Li, Pengfei Xu, Yang Zhao, Haitong Li 等ISCA 2020 · 被引用 86 次
- RePIM: Joint Exploitation of Activation and Weight Repetitions for In-ReRAM DNN AccelerationChen-Yang Tsai, Chin-Fu Nien, Tz-Ching Yu, Hung-Yu Yeh 等DAC 2021 · 被引用 22 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- YOCO: A Hybrid In-Memory Computing Architecture with 8-bit Sub-PetaOps/W In-Situ Multiply Arithmetic for Large-Scale AIZihao Xuan, Yuxuan Yang, Wei Xuan, Zijia Su 等DAC 2025 · 被引用 1 次
