MIME: adapting a single neural network for multi-task inference with memory-efficient dynamic pruning
Abhiroop Bhattacharjee, Yeshwanth Venkatesha, Abhishek Moitra, Priyadarshini Panda
摘要
Recent years have seen a paradigm shift towards multi-task learning. This calls for memory and energy-efficient solutions for inference in a multi-task scenario. We propose an algorithm-hardware co-design approach called MIME. MIME reuses the weight parameters of a trained parent task and learns task-specific threshold parameters for inference on multiple child tasks. We find that MIME results in highly memory-efficient DRAM storage of neural-network parameters for multiple tasks compared to conventional multi-task inference. In addition, MIME results in input-dependent dynamic neuronal pruning, thereby enabling energy-efficient inference with higher throughput on a systolic-array hardware. Our experiments with benchmark datasets (child tasks)- CIFAR10, CIFAR100, and Fashion-MNIST, show that MIME achieves 3.48x memory-efficiency and 2.4 - 3.1x energy-savings compared to conventional multi-task inference in Pipelined task mode.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- Proving the Lottery Ticket Hypothesis: Pruning is All You NeedEran Malach, Gilad Yehudai, Shai Shalev-Shwartz, Ohad ShamirICML 2020 · 被引用 327 次
- Dynamic Sparse Training: Find Efficient Sparse Network From Scratch With Trainable Masked LayersJunjie Liu, Zhe Xu, Runbin Shi, Ray C. C. Cheung 等ICLR 2020 · 被引用 136 次
- Model Zoo: A Growing Brain That Learns ContinuallyRahul Ramesh, Pratik ChaudhariICLR 2022 · 被引用 79 次
- PIM-Prune: Fine-Grain DCNN Pruning for Crossbar-Based Process-In-Memory ArchitectureChaoqun Chu, Yanzhi Wang, Yilong Zhao, Xiaolong Ma 等DAC 2020 · 被引用 64 次
- Learn-to-Share: A Hardware-friendly Transfer Learning Framework Exploiting Computation and Parameter SharingCheng Fu, Hanxian Huang, Xinyun Chen, Yuandong Tian 等ICML 2021 · 被引用 28 次
相关 Paper
- Algorithm/architecture co-design for energy-efficient acceleration of multi-task DNNJaekang Shin, Seungkyu Choi, Jongwoo Ra, Lee-Sup KimDAC 2022 · 被引用 4 次
- Controllable Dynamic Multi-Task ArchitecturesDripta S. Raychaudhuri, Yumin Suh, Samuel Schulter, Xiang Yu 等CVPR 2022 · 被引用 24 次
- Adyna: Accelerating Dynamic Neural Networks with Adaptive SchedulingZhiyao Li, Bohan Yang, Jiaxiang Li, Taijie Chen 等HPCA 2025 · 被引用 2 次
- EdgeMM: Multi-Core CPU with Heterogeneous AI-Extension and Activation-aware Weight Pruning for Multimodal LLMs at EdgeKangbo Bai, Le Ye, Ru Huang, Tianyu JiaDAC 2025 · 被引用 1 次
- MiniMalloc: A Lightweight Memory Allocator for Hardware-Accelerated Machine LearningMichael D. MoffittASPLOS 2023 · 被引用 5 次
