Think Fast: A Tensor Streaming Processor (TSP) for Accelerating Deep Learning Workloads
Dennis Abts, Jonathan Ross, Jonathan Sparling, Mark Wong-VanHaren, Max Baker, Tom Hawkins, Andrew Bell, John Thompson, Temesghen Kahsai, Garrin Kimmell, Jennifer Hwang, Rebekah Leslie-Hurd
摘要
In this paper, we introduce the Tensor Streaming Processor (TSP) architecture, a functionally-sliced microarchitecture with memory units interleaved with vector and matrix deep learning functional units in order to take advantage of dataflow locality of deep learning operations. The TSP is built based on two key observations: (1) machine learning workloads exhibit abundant data parallelism, which can be readily mapped to tensors in hardware, and (2) a simple and deterministic processor with producer-consumer stream programming model enables precise reasoning and control of hardware components, achieving good performance and power efficiency. The TSP is designed to exploit parallelism inherent in machine-learning workloads including instruction-level, memory concurrency, data and model parallelism, while guaranteeing determinism by eliminating all reactive elements in the hardware (e.g. arbiters, and caches). Early ResNet50 image classification results demonstrate 20.4K processed images per second (IPS) with a batch-size of one— a improvement compared to other modern GPUs and accelerators [44]. Our first ASIC implementation of the TSP architecture yields a computational density of more than 1 TeraOp/s per square mm of silicon for its mm 14nm chip operating at a nominal clock frequency of 900 MHz. The TSP demonstrates a novel hardware-software approach to achieve fast, yet predictable, performance on machine-learning workloads within a desired power envelope.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- ARK: Fully Homomorphic Encryption Accelerator with Runtime Data Generation and Inter-Operation Key ReuseJongmin Kim, Gwangho Lee, Sangpyo Kim, Gina Sohn 等MICRO 2022 · 被引用 160 次
- Not All GPUs Are Created Equal: Characterizing Variability in Large-Scale, Accelerator-Rich SystemsPrasoon Sinha, Akhil Guliani, Rutwik Jain, Brandon Tran 等SC 2022 · 被引用 31 次
- Manticore: Hardware-Accelerated RTL Simulation with Static Bulk-Synchronous ParallelismMahyar Emami, Sahand Kashani, Keisuke Kamahori, Mohammad Sepehr Pourghannad 等ASPLOS 2023 · 被引用 16 次
- Be Like Water: Adaptive Floating Point for Machine LearningThomas Y. Yeh, Max Sterner, Zerlina Lai, Brandon Chuang 等ICML 2022 · 被引用 11 次
- Cocco: Hardware-Mapping Co-Exploration towards Memory Capacity-Communication OptimizationZhanhong Tan, Zijian Zhu, Kaisheng MaASPLOS 2024 · 被引用 9 次
相关 Paper
- A software-defined tensor streaming multiprocessor for large-scale machine learningDennis Abts, Garrin Kimmell, Andrew C. Ling, John Kim 等ISCA 2022 · 被引用 46 次
- Streaming Tensor Programs: A Streaming Abstraction for Dynamic ParallelismGina Sohn, Genghan Zhang, Konstantin Hoßfeld, Jungwoo Kim 等ASPLOS 2026 · 被引用 1 次
- Tensor processing primitives: a programming abstraction for efficiency and portability in deep learning workloadsEvangelos Georganas, Dhiraj D. Kalamkar, Sasikanth Avancha, Menachem Adelman 等SC 2021 · 被引用 2 次
- StreamTensor: Make Tensors Stream in Dataflow Accelerators for LLMsHanchen Ye, Deming ChenMICRO 2025 · 被引用 5 次
- Exploring the Performance Improvement of Tensor Processing Engines through Transformation in the Bit-weight Dimension of MACsQizhe Wu, Huawen Liang, Yuchen Gui, Zhichen Zeng 等HPCA 2025 · 被引用 2 次
