Capstan: A Vector RDA for Sparsity
Alexander Rucker, Matthew Vilim, Tian Zhao, Yaqi Zhang, Raghu Prabhakar, Kunle Olukotun
Abstract
This paper proposes Capstan: a scalable, parallel-patterns-based, reconfigurable dataflow accelerator (RDA) for sparse and dense tensor applications. Instead of designing for one application, we start with common sparse data formats, each of which supports multiple applications. Using a declarative programming model, Capstan supports application-independent sparse iteration and memory primitives that can be mapped to vectorized, high-performance hardware. We optimize random-access sparse memories with configurable out-oforder execution to increase SRAM random-access throughput from 32% to 80%.
For a variety of sparse applications, Capstan with DDR4 memory is 18× faster than a multi-core CPU baseline, while Capstan with HBM2 memory is 16× faster than an Nvidia V100 GPU. For sparse applications that can be mapped to Plasticine, a recent dense RDA, Capstan is 7.6× to 365× faster and only 16% larger.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers18
- Taurus: a data plane architecture for per-packet MLTushar Swamy, Alexander Rucker, Muhammad Shahbaz, Ishan Gaur et al.ASPLOS 2022 · 94 citations
- The Sparse Abstract MachineOlivia Hsu, Maxwell Strange, Ritvik Sharma, Jaeyeon Won et al.ASPLOS 2023 · 37 citations
- SOFA: A Compute-Memory Optimized Sparsity Accelerator via Cross-Stage Coordinated TilingHuizheng Wang, Jiahao Fang, Xinru Tang, Zhiheng Yue et al.MICRO 2024 · 31 citations
- FEASTA: A Flexible and Efficient Accelerator for Sparse Tensor Algebra in Machine LearningKai Zhong, Zhenhua Zhu, Guohao Dai, Hongyi Wang et al.ASPLOS 2024 · 16 citations
- Mosaic: An Interoperable Compiler for Tensor AlgebraManya Bansal, Olivia Hsu, Kunle Olukotun, Fredrik KjolstadPLDI 2023 · 16 citations
Builds on5
- SpArch: Efficient Architecture for Sparse Matrix MultiplicationZhekai Zhang, Hanrui Wang, Song Han, William J. DallyHPCA 2020 · 280 citations
- MatRaptor: A Sparse-Sparse Matrix Multiplication Accelerator Based on Row-Wise ProductNitish Kumar Srivastava, Hanchen Jin, Jie Liu, David H. Albonesi et al.MICRO 2020 · 223 citations
- PolyGraph: Exposing the Value of Flexibility for Graph Processing AcceleratorsVidushi Dadu, Sihao Liu, Tony NowatzkiISCA 2021 · 60 citations
- SARA: Scaling a Reconfigurable Dataflow AcceleratorYaqi Zhang, Nathan Zhang, Tian Zhao, Matt Vilim et al.ISCA 2021 · 58 citations
- Gorgon: Accelerating Machine Learning from Relational DataMatthew Vilim, Alexander Rucker, Yaqi Zhang, Sophia Liu et al.ISCA 2020 · 25 citations
Related papers
- Harmonia: A Unified Hierarchical Scheduling Framework for Sparse Matrix MultiplicationJingkui Yang, Fangxin Liu, Xin Ju, Ning Yang et al.ISCA 2026
- ALRESCHA: A Lightweight Reconfigurable Sparse-Computation AcceleratorBahar Asgari, Ramyad Hadidi, Tushar Krishna, Hyesoon Kim et al.HPCA 2020 · 63 citations
- Sparsepipe: Sparse Inter-operator Dataflow Architecture with Cross-Iteration ReuseYunan Zhang, Po-An Tsai, Hung-Wei TsengMICRO 2024 · 2 citations
- Tensaurus: A Versatile Accelerator for Mixed Sparse-Dense Tensor ComputationsNitish Kumar Srivastava, Hanchen Jin, Shaden Smith, Hongbo Rong et al.HPCA 2020 · 121 citations
- Spada: Accelerating Sparse Matrix Multiplication with Adaptive DataflowZhiyao Li, Jiaxiang Li, Taijie Chen, Dimin Niu et al.ASPLOS 2023 · 59 citations
