ModularEvo: Evolving Multi-Task Models via Neural Network Modularization and Composition
Wenrui Long, Binhang Qi, Hailong Sun, Zongzhen Yang, Ruobing Zhao, Xiang Gao
摘要
Training a general multi-task deep neural network (DNN) model, such as a large language model, and deploying it across diverse downstream tasks has become a common practice. In long-term deployment scenarios, downstream tasks can change over time, such as new data distributions and requirements, leading to the fine-tuning of the model accordingly, i.e., evolving the model. However, traditional full-parameter fine-tuning methods adapt the model to individual tasks, resulting in degradation of the original knowledge. Although parameter-efficient fine-tuning methods could mitigate this problem, they still isolate new knowledge in external, separate parameters. As a result, the base model gains little cumulative benefit from downstream updates. These limitations stem from the indiscriminate model deployment and fine-tuning. Inspired by modular design principles in software engineering, we propose ModularEvo, a framework that enables on-demand deployment and co-evolution of multi-task models and modules across diverse downstream tasks. ModularEvo first decomposes the model into task-specific modules, each retaining a subset of relevant weights and functionality. These modules, instead of the entire model, are deployed on downstream tasks on demand. During long-term deployment, each module is independently optimized to adjust to the change of the corresponding task. Unlike conventional fine-tuning methods, ModularEvo applies modular fine-tuning to update only the task-relevant weights within modules. Furthermore, new knowledge acquired by modules is periodically integrated into the model through model merging, enabling the co-evolution of both the model and modules. We evaluate ModularEvo through extensive experiments on three Transformer models and six downstream tasks involving both classification and generation tasks. Results demonstrate the effectiveness of ModularEvo in model performance and inference efficiency in evolution scenarios. Compared to state-of-the-art baselines, ModularEvo achieves an absolute performance gain of 2.34% in multi-round evolution scenarios and a 2.22 times speedup in inference. In all, our work provides a new paradigm for the reuse and evolution of DNN models in the development of intelligent software applications.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Modularizing while Training: A New Paradigm for Modularizing DNN ModelsBinhang Qi, Hailong Sun, Hongyu Zhang, Ruobing Zhao 等ICSE 2024 · 被引用 3 次
- MoDULA: Mixture of Domain-Specific and Universal LoRA for Multi-Task LearningYufei Ma, Zihan Liang, Huangyu Dai, Ben Chen 等EMNLP 2024 · 被引用 4 次
- On decomposing a deep neural network into modulesRangeet Pan, Hridesh RajanFSE 2020 · 被引用 38 次
- Adapt Once, Thrive with Updates: Transferable Parameter-Efficient Fine-Tuning on Evolving Base ModelsNaibin Gu, Peng Fu, Xiyu Liu, Ke Ma 等ACL 2025
- Evolving Subnetwork Training for Large Language ModelsHanqi Li, Lu Chen, Da Ma, Zijian Wu 等ICML 2024 · 被引用 2 次
