Model LEGO: Creating Models Like Disassembling and Assembling Building Blocks
Jiacong Hu, Jing Gao, Jingwen Ye, Yang Gao, Xingen Wang, Zunlei Feng, Mingli Song
摘要
With the rapid development of deep learning, the increasing complexity and scale of parameters make training a new model increasingly resource-intensive. In this paper, we start from the classic convolutional neural network (CNN) and explore a paradigm that does not require training to obtain new models. Similar to the birth of CNN inspired by receptive fields in the biological visual system, we draw inspiration from the information subsystem pathways in the biological visual system and propose Model Disassembling and Assembling (MDA). During model disassembling, we introduce the concept of relative contribution and propose a component locating technique to extract task-aware components from trained CNN classifiers. For model assembling, we present the alignment padding strategy and parameter scaling strategy to construct a new model tailored for a specific task, utilizing the disassembled task-aware components. The entire process is akin to playing with LEGO bricks, enabling arbitrary assembly of new models, and providing a novel perspective for model creation and reuse. Extensive experiments showcase that task-aware components disassembled from CNN classifiers or new models assembled using these components closely match or even surpass the performance of the baseline, demonstrating its promising results for model reuse. Furthermore, MDA exhibits diverse potential applications, with comprehensive experiments exploring model decision route analysis, model compression, knowledge distillation, and more. The code is available at https://github.com/jiaconghu/Model-LEGO.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper8
- Deep Model ReassemblyXingyi Yang, Daquan Zhou, Songhua Liu, Jingwen Ye 等NeurIPS 2022 · 被引用 162 次
- Architecture Disentanglement for Deep Neural NetworksJie Hu, Liujuan Cao, Tong Tong, Qixiang Ye 等ICCV 2021 · 被引用 21 次
- Model Doctor: A Simple Gradient Aggregation Strategy for Diagnosing and Treating CNN ClassifiersZunlei Feng, Jiacong Hu, Sai Wu, Xiaotian Yu 等AAAI 2022 · 被引用 16 次
- HRank: Filter Pruning Using High-Rank Feature MapMingbao Lin, Rongrong Ji, Yan Wang, Yichen Zhang 等CVPR 2020
- Neural Response Interpretation Through the Lens of Critical PathwaysAshkan Khakzar, Soroosh Baselizadeh, Saurabh Khanduja, Christian Rupprecht 等CVPR 2021
相关 Paper
- Decomposing Convolutional Neural Networks into Reusable and Replaceable ModulesRangeet Pan, Hridesh RajanICSE 2022 · 被引用 30 次
- Reusing Deep Neural Network Models through Model Re-engineeringBinhang Qi, Hailong Sun, Xiang Gao, Hongyu Zhang 等ICSE 2023 · 被引用 16 次
- Patching Weak Convolutional Neural Network Models through Modularization and CompositionBinhang Qi, Hailong Sun, Xiang Gao, Hongyu ZhangASE 2022 · 被引用 13 次
- Modularizing while Training: A New Paradigm for Modularizing DNN ModelsBinhang Qi, Hailong Sun, Hongyu Zhang, Ruobing Zhao 等ICSE 2024 · 被引用 3 次
- Building LLMs Like LEGO: Two-dimensional Architecture Reassembly of Large Language ModelsXingyu Wu, Yu Zhou, Kay Chen TanACL 2026
