ModularEvo: Evolving Multi-Task Models via Neural Network Modularization and Composition
Wenrui Long, Binhang Qi, Hailong Sun, Zongzhen Yang, Ruobing Zhao, Xiang Gao
Abstract
Training a general multi-task deep neural network (DNN) model, such as a large language model, and deploying it across diverse downstream tasks has become a common practice. In long-term deployment scenarios, downstream tasks can change over time, such as new data distributions and requirements, leading to the fine-tuning of the model accordingly, i.e., evolving the model. However, traditional full-parameter fine-tuning methods adapt the model to individual tasks, resulting in degradation of the original knowledge. Although parameter-efficient fine-tuning methods could mitigate this problem, they still isolate new knowledge in external, separate parameters. As a result, the base model gains little cumulative benefit from downstream updates. These limitations stem from the indiscriminate model deployment and fine-tuning. Inspired by modular design principles in software engineering, we propose ModularEvo, a framework that enables on-demand deployment and co-evolution of multi-task models and modules across diverse downstream tasks. ModularEvo first decomposes the model into task-specific modules, each retaining a subset of relevant weights and functionality. These modules, instead of the entire model, are deployed on downstream tasks on demand. During long-term deployment, each module is independently optimized to adjust to the change of the corresponding task. Unlike conventional fine-tuning methods, ModularEvo applies modular fine-tuning to update only the task-relevant weights within modules. Furthermore, new knowledge acquired by modules is periodically integrated into the model through model merging, enabling the co-evolution of both the model and modules. We evaluate ModularEvo through extensive experiments on three Transformer models and six downstream tasks involving both classification and generation tasks. Results demonstrate the effectiveness of ModularEvo in model performance and inference efficiency in evolution scenarios. Compared to state-of-the-art baselines, ModularEvo achieves an absolute performance gain of 2.34% in multi-round evolution scenarios and a 2.22 times speedup in inference. In all, our work provides a new paradigm for the reuse and evolution of DNN models in the development of intelligent software applications.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get e4c5bb69-8eb3-4777-b89f-7a7cdcb6ddd4Related papers
- Modularizing while Training: A New Paradigm for Modularizing DNN ModelsBinhang Qi, Hailong Sun, Hongyu Zhang, Ruobing Zhao et al.ICSE 2024 · 3 citations
- MoDULA: Mixture of Domain-Specific and Universal LoRA for Multi-Task LearningYufei Ma, Zihan Liang, Huangyu Dai, Ben Chen et al.EMNLP 2024 · 4 citations
- On decomposing a deep neural network into modulesRangeet Pan, Hridesh RajanFSE 2020 · 38 citations
- Adapt Once, Thrive with Updates: Transferable Parameter-Efficient Fine-Tuning on Evolving Base ModelsNaibin Gu, Peng Fu, Xiyu Liu, Ke Ma et al.ACL 2025
- Evolving Subnetwork Training for Large Language ModelsHanqi Li, Lu Chen, Da Ma, Zijian Wu et al.ICML 2024 · 2 citations
