TensorPrism: Rethinking Sparse High-Order Tensor Acceleration via Co-Occurrence Graph
Fangzhou Ye, Shilin Tian, Amir Ghazizadeh Ahsaei, Hao Zheng
Abstract
Sparse high-order tensors are a key computational primitive across diverse domains, including large language models, scientific computing, recommendation systems, and multi-dimensional signal processing. Existing work primarily relies on tensor contraction to unfold high-order tensors into twodimensional sparse-dense matrix multiplications (SpMM), enabling direct reuse of well-studied SpMM optimization techniques. A complementary line of research adopts tensor-native approaches, which operate directly on the multi-dimensional tensor structure to improve data locality. However, both unfoldingbased and tensor-native techniques are inefficient to fully exploit data reuse in sparse high-order tensor computations, due to (1) the expansion of matrix dimensions and (2) missed reuse opportunities across different tensor modes. In this paper, we posit that matricization dismantles high-dimensional data reuse, erasing the correlations among nonzero elements across multiple tensor modes. We propose TensorPrism, a novel acceleration framework for sparse high-order tensor computation based on a co-occurrence graph abstraction. The central idea is to transform a high-order tensor into a co-occurrence graph that captures nonzero correlations across all tensor dimensions. Building on this abstraction, TensorPrism introduces three key designs. First, we formulate a co-occurrence graph representation that redefines dataflow and tiling to improve data reuse. Second, we introduce a new dataflow strategy that enhances reuse opportunities across tensor modes. Finally, we provide an efficient accelerator design tailored to the graph-based computation. Our evaluation shows that TensorPrism delivers performance speedups of , and over state-of-the-art designs SPADE [1], HotTiles [2], GSpTC [3], and TCP [4], respectively.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get eefa6984-4ed4-4d1d-997c-35923e47d348Related papers
- Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor CoresHaisha Zhao, San Li, Jiaheng Wang, Chunbao Zhou et al.PPoPP 2025 · 18 citations
- TensorIR: An Abstraction for Automatic Tensorized Program OptimizationSiyuan Feng, Bohan Hou, Hongyi Jin, Wuwei Lin et al.ASPLOS 2023 · 80 citations
- Exploiting Efficient Mapping and Pipelined Execution for Accelerating SpMV on Tensor CoresKaige Zhang, Hailong Yang, Xin You, Tianyu Feng et al.PPoPP 2026
- Rethinking Tiling and Dataflow for SpMM Acceleration: A Graph Transformation FrameworkAmir Ghazizadeh Ahsaei, Lingxiang Yin, Shilin Tian, Fangzhou Ye et al.MICRO 2025 · 4 citations
- High Performance Unstructured SpMM Computation Using Tensor CoresPatrik Okanovic, Grzegorz Kwasniewski, Paolo Sylos Labini, Maciej Besta et al.SC 2024 · 15 citations
