Tensor processing primitives: a programming abstraction for efficiency and portability in deep learning workloads
Evangelos Georganas, Dhiraj D. Kalamkar, Sasikanth Avancha, Menachem Adelman, Cristina Anderson, Alexander Breuer, Jeremy Bruestle, Narendra Chaudhary, Abhisek Kundu, Denise Kutnick, Frank Laub, Md. Vasimuddin
摘要
During the past decade, novel Deep Learning (DL) algorithms, workloads and hardware have been developed to tackle a wide range of problems. Despite the advances in workload and hardware ecosystems, the programming methodology of DL systems is stagnant. DL workloads leverage either highly-optimized, yet platform-specific and inflexible kernels from DL libraries, or in the case of novel operators, reference implementations are built via DL framework primitives with underwhelming performance. This work introduces the Tensor Processing Primitives (TPP), a programming abstraction striving for efficient, portable implementation of DL workloads with high-productivity. TPPs define a compact, yet versatile set of 2D-tensor operators (or a virtual Tensor ISA), which subsequently can be utilized as building-blocks to construct complex operators on high-dimensional tensors. The TPP specification is platform-agnostic, thus code expressed via TPPs is portable, whereas the TPP implementation is highly-optimized and platform-specific. We demonstrate the efficacy and viability of our approach using standalone kernels and end-to-end DL & HPC workloads expressed entirely via TPPs that outperform state-of-the-art implementations on multiple platforms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- ROLLER: Fast and Efficient Tensor Compilation for Deep LearningHongyu Zhu, Ruofan Wu, Yijia Diao, Shanbin Ke 等OSDI 2022 · 被引用 84 次
- Optimizing Direct Convolutions on ARM Multi-CoresPengyu Wang, Weiling Yang, Jianbin Fang, Dezun Dong 等SC 2023 · 被引用 6 次
- DECA: A Near-Core LLM Decompression Accelerator Grounded on a 3D Roofline ModelGerasimos Gerogiannis, Stijn Eyerman, Evangelos Georganas, Wim Heirman 等MICRO 2025 · 被引用 5 次
它引用的顶会 Paper2
- Ansor: Generating High-Performance Tensor Programs for Deep LearningLianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu 等OSDI 2020 · 被引用 551 次
- Optimizing deep learning recommender systems training on CPU cluster architecturesDhiraj D. Kalamkar, Evangelos Georganas, Sudarshan Srinivasan, Jianping Chen 等SC 2020 · 被引用 41 次
相关 Paper
- TensorIR: An Abstraction for Automatic Tensorized Program OptimizationSiyuan Feng, Bohan Hou, Hongyi Jin, Wuwei Lin 等ASPLOS 2023 · 被引用 80 次
- DeepCuts: a deep learning optimization framework for versatile GPU workloadsWookeun Jung, Thanh Tuan Dao, Jaejin LeePLDI 2021 · 被引用 27 次
- TGraph: A Tensor-centric Graph Processing FrameworkYongliang Zhang, Yuanyuan Zhu, Hao Zhang, Congli Gao 等SIGMOD 2025 · 被引用 1 次
- TQEx: Tensor-based Query Engine Enhanced by Bridging the GapHaitao Zhang, Ran Pang, Yuanyuan Zhu, Hao Zhang 等SIGMOD 2026
- TAIDL: Tensor Accelerator ISA Definition Language with Auto-generation of Scalable Test OraclesDevansh Jain, Marco Frigo, Jai Arora, Akash Pardeshi 等MICRO 2025 · 被引用 1 次
