Tandem Processor: Grappling with Emerging Operators in Neural Networks
Soroush Ghodrati, Sean Kinzer, Hanyang Xu, Rohan Mahapatra, Yoonsung Kim, Byung Hoon Ahn, Dong Kai Wang, Lavanya Karthikeyan, Amir Yazdanbakhsh, Jongse Park, Nam Sung Kim, Hadi Esmaeilzadeh
摘要
With the ever increasing prevalence of neural networks and the upheaval from the language models, it is time to rethink neural acceleration. Up to this point, the broader research community, including ourselves, has disproportionately focused on GEneral Matrix Multiplication (GEMM) operations. The supporting argument was that the large majority of the neural operations are GEMM. This argument guided the research in Neural Processing Units (NPUs) for the last decade. However, scant attention was paid to non-GEMM operations and they are rather overlooked. As deep learning evolved and progressed, these operations have grown in diversity and also large variety of structural patterns have emerged that interweave them with the GEMM operations. However, conventional NPU designs have taken rather simplistic approaches by supporting these operations through either a number of dedicated blocks or fall back to general-purpose processors.
This work sets out to challenge the conventional wisdom in neural accelerator design and explore the architecture of an on-chip companion, dubbed Tandem Processor, that complements the rather optimized GEMM unit in neural accelerators. This processor needs to be specialized to keep up with the GEMM unit; and yet needs to be programmable to address the (1) structural and (2) operational variations. To strike a balance between specialization and programmability, on the one hand, we specialize its memory access logic with a novel ISA/microarchitecture that alleviates the register file and its associated load/store operations. On the other hand, the calculations of the non-GEMM layers are only supported through primitive arithmetic/logic vector operations. Therefore, programmability is offered at the mathematical level. The enhancements due to the specialization of the memory access logic in the Tandem Processor and its tight integration with the GEMM unit sustain the throughput and the utilization of the neural accelerator. Comprehensive evaluations of the proposed design based on the end-to-end execution of seven diverse DNNs including emerging language models show significant performance improvements and energy reduction enabled by leveraging the Tandem Processor. We provide the RTL code that is synthesizable both for FPGA and ASIC implementations in addition to the associated compiler as part of the open-source GeneSys project (https://actlab-genesys. github.io/). We also present the chip floorplan and post-layout analysis. This work is the result of 10 years of effort in building real NPUs that support end-to-end neural network execution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- PICACHU: Plug-In CGRA Handling Upcoming Nonlinear Operations in LLMsJiajun Qin, Tianhua Xia, Cheng Tan, Jeff Zhang 等ASPLOS 2025 · 被引用 17 次
- In-Storage Acceleration of Retrieval Augmented Generation as a ServiceRohan Mahapatra, Harsha Santhanam, Christopher Priebe, Hanyang Xu 等ISCA 2025 · 被引用 9 次
- In-Storage Domain-Specific Acceleration for Serverless ComputingRohan Mahapatra, Soroush Ghodrati, Byung Hoon Ahn, Sean Kinzer 等ASPLOS 2024 · 被引用 7 次
- DECA: A Near-Core LLM Decompression Accelerator Grounded on a 3D Roofline ModelGerasimos Gerogiannis, Stijn Eyerman, Evangelos Georganas, Wim Heirman 等MICRO 2025 · 被引用 5 次
- Integrated Hardware Architecture and Device Placement SearchIrene Wang, Jakub Tarnawski, Amar Phanishayee, Divya MahajanICML 2024 · 被引用 4 次
它引用的顶会 Paper12
- SIGMA: A Sparse and Irregular GEMM Accelerator with Flexible Interconnects for DNN TrainingEric Qin, Ananda Samajdar, Hyoukjun Kwon, Vineet Nadella 等HPCA 2020 · 被引用 490 次
- I-BERT: Integer-only BERT QuantizationSehoon Kim, Amir Gholami, Zhewei Yao, Michael W. Mahoney 等ICML 2021 · 被引用 439 次
- A3: Accelerating Attention Mechanisms in Neural Networks with ApproximationTae Jun Ham, Sungjun Jung, Seonghak Kim, Young H. Oh 等HPCA 2020 · 被引用 241 次
- HAWQ-V3: Dyadic Neural Network QuantizationZhewei Yao, Zhen Dong, Zhangcheng Zheng, Amir Gholami 等ICML 2021 · 被引用 240 次
- MatRaptor: A Sparse-Sparse Matrix Multiplication Accelerator Based on Row-Wise ProductNitish Kumar Srivastava, Hanchen Jin, Jie Liu, David H. Albonesi 等MICRO 2020 · 被引用 223 次
相关 Paper
- NeuMMU: Architectural Support for Efficient Address Translations in Neural Processing UnitsBongjoon Hyun, Youngeun Kwon, Yujeong Choi, John Kim 等ASPLOS 2020 · 被引用 29 次
- Accelerating applications using edge tensor processing unitsKuan-Chieh Hsu, Hung-Wei TsengSC 2021 · 被引用 33 次
- DataMaestro: A Versatile and Efficient Data Streaming Engine Bringing Decoupled Memory Access To Dataflow AcceleratorsXiaoling Yi, Yunhao Deng, Ryan Antonio, Fanchen Kong 等DAC 2025 · 被引用 4 次
- NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM InferencingGuseul Heo, Sangyeop Lee, Jaehong Cho, Hyunmin Choi 等ASPLOS 2024 · 被引用 121 次
- GPNPU: Enabling Efficient Hardware-Based Direct Convolution with Multi-Precision Support in GPU Tensor CoresZhuoran Song, Jianfei Wang, Tianjian Li, Li Jiang 等DAC 2020 · 被引用 12 次
