MINGLE: Mixture of Null-Space Gated Low-Rank Experts for Test-Time Continual Model Merging
Zihuan Qiu, Yi Xu, Chiyuan He, Fanman Meng, Linfeng Xu, Qingbo Wu, Hongliang Li
摘要
Continual model merging integrates independently fine-tuned models sequentially without access to the original training data, offering a scalable and efficient solution for continual learning. However, existing methods face two critical challenges: parameter interference among tasks, which leads to catastrophic forgetting, and limited adaptability to evolving test distributions. To address these issues, we introduce the task of Test-Time Continual Model Merging (TTCMM), which leverages a small set of unlabeled test samples during inference to alleviate parameter conflicts and handle distribution shifts. We propose MINGLE, a novel framework for TTCMM. MINGLE employs a mixture-of-experts architecture with parameter-efficient, low-rank experts, which enhances adaptability to evolving test distributions while dynamically merging models to mitigate conflicts. To further reduce forgetting, we propose Null-Space Constrained Gating, which restricts gating updates to subspaces orthogonal to prior task representations, thereby suppressing activations on old tasks and preserving past knowledge. We further introduce an Adaptive Relaxation Strategy that adjusts constraint strength dynamically based on interference signals observed during test-time adaptation, striking a balance between stability and adaptability. Extensive experiments on standard continual merging benchmarks demonstrate that MINGLE achieves robust generalization, significantly reduces forgetting, and consistently surpasses previous state-of-the-art methods by 7-9% on average across diverse task orders. Our code is available at: https://github.com/zihuanqiu/MINGLE
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- ACE-Merging: Data-Free Model Merging with Adaptive Covariance EstimationBo Xu, Haotian Wu, Hehai Lin, Weiquan Huang 等CVPR 2026 · 被引用 5 次
- Null-Space Filtering for Data-Free Continual Model Merging: Preserving Stability, Promoting PlasticityZihuan Qiu, Lei Wang, Yang Cao, Runtong ZHANG 等ICLR 2026 · 被引用 4 次
- Sparsity Curse: Understanding RLVR Model Parameter Space from Model MergingChenrui Wu, Zexi Li, Jiajun Bu, Jiangchuan Liu 等KDD 2026 · 被引用 2 次
- Merge to Remember: Sharpness-Aware Isotropic Merging for Continual LearningQun Yang, Enneng Yang, Wei Chen, Li Shen 等ICML 2026
- Efficient Bilevel Optimization for CKA-Guided MoE UpcyclingZhiyuan Yu, Enneng Yang, Hao Jiang, Guojie Zhu 等ICML 2026
它引用的顶会 Paper47
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging FaceYongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li 等NeurIPS 2023 · 被引用 1,778 次
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen 等ICLR 2021 · 被引用 1,731 次
- Test-Time Training with Self-Supervision for Generalization under Distribution ShiftsYu Sun, Xiaolong Wang, Zhuang Liu, John Miller 等ICML 2020 · 被引用 1,220 次
相关 Paper
- Train with Perturbation, Infer after Merging: A Two-Stage Framework for Continual LearningHaomiao Qiu, Miao Zhang, Ziyue Qiao, Liqiang NieNeurIPS 2025 · 被引用 8 次
- Theory on Mixture-of-Experts in Continual LearningHongbo Li, Sen Lin, Lingjie Duan, Yingbin Liang 等ICLR 2025
- BECoTTA: Input-dependent Online Blending of Experts for Continual Test-time AdaptationDaeun Lee, Jaehong Yoon, Sung Ju HwangICML 2024 · 被引用 27 次
- Mixture of Prototypes for Test-time Adaptive SegmentationGuangrui Li, Zhengyu Zhu, Yongxin GeCVPR 2026 · 被引用 1 次
- From Experts to Bases: Orthogonal Subspace Mixture for Continual Multimodal Instruction TuningPei Chen, Xilai Wang, Qixu Shi, Zejian Li 等ACL 2026
