Sparta: high-performance, element-wise sparse tensor contraction on heterogeneous memory
Jiawen Liu, Jie Ren, Roberto Gioiosa, Dong Li, Jiajia Li
摘要
Sparse tensor contractions appear commonly in many applications. Efficiently computing a two sparse tensor product is challenging: It not only inherits the challenges from common sparse matrix-matrix multiplication (SpGEMM), i.e., indirect memory access and unknown output size before computation, but also raises new challenges because of high dimensionality of tensors, expensive multi-dimensional index search, and massive intermediate and output data. To address the above challenges, we introduce three optimization techniques by using multi-dimensional, efficient hashtable representation for the accumulator and larger input tensor, and all-stage parallelization. Evaluating with 15 datasets, we show that Sparta brings 28 -- 576× speedup over the traditional sparse tensor contraction with sparse accumulator. With our proposed algorithm- and memory heterogeneity-aware data management, Sparta brings extra performance improvement on the heterogeneous memory with DRAM and Intel Optane DC Persistent Memory Module (PMM) over a state-of-the-art software-based data management solution, a hardware-based data management solution, and PMM-only by 30.7% (up to 98.5%), 10.7% (up to 28.3%) and 17% (up to 65.1%) respectively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- FlexMem: Adaptive Page Profiling and Migration for Tiered MemoryDong Xu, Junhee Ryu, Kwangsik Shin, Pengfei Su 等USENIX ATC 2024 · 被引用 41 次
- Efficient Quantized Sparse Matrix Operations on Tensor CoresShigang Li, Kazuki Osawa, Torsten HoeflerSC 2022 · 被引用 27 次
- Merchandiser: Data Placement on Heterogeneous Memory for Task-Parallel HPC Applications with Load-Balance AwarenessZhen Xie, Jie Liu, Jiajia Li, Dong LiPPoPP 2023 · 被引用 18 次
- SparseAuto: An Auto-scheduler for Sparse Tensor Computations using Recursive Loop Nest RestructuringAdhitha Dias, Logan Anderson, Kirshanthan Sundararajah, Artem Pelenitsyn 等OOPSLA 2024 · 被引用 6 次
- Machine Learning-Guided Memory Optimization for DLRM Inference on Tiered MemoryJie Ren, Bin Ma, Shuangyan Yang, Benjamin Francis 等HPCA 2025 · 被引用 6 次
它引用的顶会 Paper6
- An Empirical Guide to the Behavior and Use of Scalable Persistent MemoryJian Yang, Juno Kim, Morteza Hoseinzadeh, Joseph Izraelevitz 等FAST 2020 · 被引用 470 次
- HM-ANN: Efficient Billion-Point Nearest Neighbor Search on Heterogeneous MemoryJie Ren, Minjia Zhang, Dong LiNeurIPS 2020 · 被引用 136 次
- Sentinel: Efficient Tensor Migration and Allocation on Heterogeneous Memory Systems for Deep LearningJie Ren, Jiaolin Luo, Kai Wu, Minjia Zhang 等HPCA 2021 · 被引用 62 次
- ArchTM: Architecture-Aware, High Performance Transaction for Persistent MemoryKai Wu, Jie Ren, Ivy Bo Peng, Dong LiFAST 2021 · 被引用 31 次
- Distributed-memory DMRG via sparse and dense parallel tensor contractionsRyan Levy, Edgar Solomonik, Bryan K. ClarkSC 2020 · 被引用 11 次
相关 Paper
- Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor CoresHaisha Zhao, San Li, Jiaheng Wang, Chunbao Zhou 等PPoPP 2025 · 被引用 18 次
- SpArch: Efficient Architecture for Sparse Matrix MultiplicationZhekai Zhang, Hanrui Wang, Song Han, William J. DallyHPCA 2020 · 被引用 280 次
- FaSTCC: Fast Sparse Tensor Contractions on CPUsSaurabh Raje, Hunter McCoy, Atanas Rountev, Prashant Pandey 等SC 2025 · 被引用 2 次
- A Probabilistic Perspective on Tiling Sparse Tensor AlgebraRitvik Sharma, Zi Yu Xue, Nathan Zhang, Rubens Lacouture 等MICRO 2025 · 被引用 1 次
- Exploiting Hierarchical Parallelism and Reusability in Tensor Kernel Processing on Heterogeneous HPC SystemsYuedan Chen, Guoqing Xiao, M. Tamer Özsu, Zhuo Tang 等ICDE 2022 · 被引用 7 次
