DNN Modularization via Activation-Driven Training
Tuan Ngo, Abid Hassan, Saad Shafiq, Nenad Medvidović
摘要
Deep Neural Networks (DNNs) tend to accrue technical debt and suffer from significant retraining costs when adapting to evolving requirements. Modularizing DNNs offers the promise of improving their reusability. Previous work has proposed techniques to decompose DNN models into modules both during and after training. However, these strategies yield several shortcomings, including significant weight overlaps and accuracy losses across modules, restricted focus on convolutional layers only, and added complexity and training time by introducing auxiliary masks to control modularity. In this work, we propose MODA, an activation-driven modular training approach. MODA promotes inherent modularity within a DNN model by directly regulating the activation outputs of its layers based on three modular objectives: intra-class affinity, inter-class dispersion, and compactness. MODA is evaluated using three well-known DNN models and five datasets with varying sizes. This evaluation indicates that, compared to the existing state-of-the-art, using MODA yields several advantages: (1) MODA accomplishes modularization with 22% less training time; (2) the resultant modules generated by MODA comprise up to 24x fewer weights and 37x less weight overlap while (3) preserving the original model’s accuracy without additional fine-tuning; in module replacement scenarios, (4) MODA improves the accuracy of a target class by 12% on average while ensuring minimal impact on the accuracy of other classes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine 等AAAI 2020 · 被引用 1,361 次
- DeepGini: prioritizing massive tests to enhance the robustness of deep neural networksYang Feng, Qingkai Shi, Xinyu Gao, Jun Wan 等ISSTA 2020 · 被引用 206 次
- On decomposing a deep neural network into modulesRangeet Pan, Hridesh RajanFSE 2020 · 被引用 38 次
- Decomposing Convolutional Neural Networks into Reusable and Replaceable ModulesRangeet Pan, Hridesh RajanICSE 2022 · 被引用 30 次
相关 Paper
- Modularizing while Training: A New Paradigm for Modularizing DNN ModelsBinhang Qi, Hailong Sun, Hongyu Zhang, Ruobing Zhao 等ICSE 2024 · 被引用 3 次
- DeepArc: Modularizing Neural Networks for the Model MaintenanceXiaoning Ren, Yun Lin, Yinxing Xue, Ruofan Liu 等ICSE 2023 · 被引用 6 次
- DIVISION: Memory Efficient Training via Dual Activation PrecisionGuanchu Wang, Zirui Liu, Zhimeng Jiang, Ninghao Liu 等ICML 2023 · 被引用 4 次
- AC/DC: Alternating Compressed/DeCompressed Training of Deep Neural NetworksAlexandra Peste, Eugenia Iofinova, Adrian Vladu, Dan AlistarhNeurIPS 2021 · 被引用 84 次
- Patching Weak Convolutional Neural Network Models through Modularization and CompositionBinhang Qi, Hailong Sun, Xiang Gao, Hongyu ZhangASE 2022 · 被引用 13 次
