Dataflow Mirroring: Architectural Support for Highly Efficient Fine-Grained Spatial Multitasking on Systolic-Array NPUs
Jounghoo Lee, Jinwoo Choi, Jaeyeon Kim, Jinho Lee, Youngsok Kim
Abstract
We present dataflow mirroring, architectural support for low-overhead fine-grained systolic array allocation which overcomes the limitations of prior coarse-grained spatial-multitasking Neural Processing Unit (NPU) architectures. The key idea of dataflow mirroring is to reverse the dataflows of co-located Neural Networks (NNs) in horizontal and/or vertical directions, allowing allocation boundaries to be set between any adjacent rows and columns of a systolic array and supporting up to four-way spatial multitasking. Our detailed experiments using MLPerf NNs and a dataflow-mirroring-augmented NPU prototype which extends Google’s TPU with dataflow mirroring shows that dataflow mirroring can significantly improve the multitasking performance by up to 46.4%.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get a3849644-a129-44c5-9e8b-228c9d272f6fCited by top-tier papers6
- DACAPO: Accelerating Continuous Learning in Autonomous Systems for Video AnalyticsYoonsung Kim, Changhun Oh, Jinwoo Hwang, Wonung Kim et al.ISCA 2024 · 13 citations
- Sparse-DySta: Sparsity-Aware Dynamic and Static Scheduling for Sparse Multi-DNN WorkloadsHongxiang Fan, Stylianos I. Venieris, Alexandros Kouris, Nicholas D. LaneMICRO 2023 · 12 citations
- AuRORA: Virtualized Accelerator Orchestration for Multi-Tenant WorkloadsSeah Kim, Jerry Zhao, Krste Asanovic, Borivoje Nikolic et al.MICRO 2023 · 10 citations
- XSched: Preemptive Scheduling for Diverse XPUsWeihang Shen, Mingcong Han, Jialong Liu, Rong Chen et al.OSDI 2025 · 9 citations
- Perceptual-Centric Image Super-Resolution using Heterogeneous Processors on Mobile DevicesKai Huang, Xiangyu Yin, Tao Gu, Wei GaoMobiCom 2024 · 6 citations
Related papers
- UPTPU: Improving Energy Efficiency of a Tensor Processing Unit through Underutilization Based Power-GatingPramesh Pandey, Noel Daniel Gundi, Koushik Chakraborty, Sanghamitra RoyDAC 2021 · 10 citations
- V10: Hardware-Assisted NPU Multi-tenancy for Improved Resource Utilization and FairnessYuqi Xue, Yiqi Liu, Lifeng Nai, Jian HuangISCA 2023 · 21 citations
- A Tensor Marshaling Unit for Sparse Tensor Algebra on General-Purpose ProcessorsMarco Siracusa, Víctor Soria Pardos, Francesco Sgherzi, Joshua Randall et al.MICRO 2023 · 11 citations
- A Multi-Neural Network Acceleration ArchitectureEunjin Baek, Dongup Kwon, Jangwoo KimISCA 2020 · 110 citations
- Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasksLingxiao Ma, Zhiqiang Xie, Zhi Yang, Jilong Xue et al.OSDI 2020 · 192 citations
