VGA: Hardware Accelerator for Scalable Long Sequence Model Inference
Seung Yul Lee, Hyunseung Lee, Jihoon Hong, SangLyul Cho, Jae W. Lee
Abstract
Effectively modeling relationships between distant elements in a long input sequence is an important task that remains challenging to this day. The state-of-the-art models for processing sequential data are self-attention-based transformer models. However, the computational complexity of self-attention is quadratic to the input sequence length, which often becomes the limiting factor in scaling the sequence length. Recently, state space model (SSM)-based global convolution models, which replace attention with convolution, have been found to be effective for modeling long sequences, with a sub-quadratic complexity using Fast Fourier Transform (FFT). However, they show sub-optimal performance on data-parallel accelerators like GPU, due to the regions of extremely low compute utilization with memory bandwidth-bound operations. To address this inefficiency, this paper proposes the Vandermonde matrix Generating Accelerator (VGA), a custom accelerator that performs FFT-based convolution in an area/power-efficient manner. VGA introduces Complex number Compute Units (CCUs) to fully utilize the high on-chip SRAM bandwidth, and parameters are generated on the fly to drastically reduce the required SRAM capacity. VGA achieves 76×(48×) higher area (power) efficiency than NVIDIA A100 GPU when executing the global convolution operator of H3, a state-of-the-art SSM-based model.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 697abea3-ecde-45e6-99fa-e2f4164d0af4Cited by top-tier papers2
- Pimba: A Processing-in-Memory Acceleration for Post-Transformer Large Language Model ServingWonung Kim, Yubin Lee, Yoonsung Kim, Jinwoo Hwang et al.MICRO 2025 · 6 citations
- Chameleon: Adaptive Caching and Scheduling for Many-Adapter LLM Inference EnvironmentsNikoleta Iliakopoulou, Jovan Stojkovic, Chloe Alverti, Tianyin Xu et al.MICRO 2025 · 3 citations
Related papers
- What Makes Convolutional Models Great on Long Sequence Modeling?Yuhong Li, Tianle Cai, Yi Zhang, Deming Chen et al.ICLR 2023 · 20 citations
- Hungry Hungry Hippos: Towards Language Modeling with State Space ModelsDaniel Y. Fu, Tri Dao, Khaled Kamal Saab, Armin W. Thomas et al.ICLR 2023 · 117 citations
- Simple Hardware-Efficient Long Convolutions for Sequence ModelingDaniel Y. Fu, Elliot L. Epstein, Eric Nguyen, Armin W. Thomas et al.ICML 2023 · 72 citations
- Short-Long Convolutions Help Hardware-Efficient Linear Attention to Focus on Long SequencesZicheng Liu, Siyuan Li, Li Wang, Zedong Wang et al.ICML 2024 · 11 citations
- Convolutional State Space Models for Long-Range Spatiotemporal ModelingJimmy T. H. Smith, Shalini De Mello, Jan Kautz, Scott W. Linderman et al.NeurIPS 2023 · 36 citations
