Federated Continual Instruction Tuning
Haiyang Guo, Fanhu Zeng, Fei Zhu, Wenzhuo Liu, Da-Han Wang, Jian Xu, Xu-Yao Zhang, Cheng-Lin Liu
Abstract
A vast amount of instruction tuning data is crucial for the impressive performance of Large Multimodal Models (LMMs), but the associated computational costs and data collection demands during supervised fine-tuning make it impractical for most researchers. Federated learning (FL) has the potential to leverage all distributed data and training resources to reduce the overhead of joint training. However, most existing methods assume a fixed number of tasks, while in real-world scenarios, clients continuously encounter new knowledge and often struggle to retain old tasks due to memory constraints. In this work, we introduce the Federated Continual Instruction Tuning (FCIT) benchmark to model this real-world challenge. Our benchmark includes two realistic scenarios, encompassing four different settings and twelve carefully curated instruction tuning datasets. To address the challenges posed by FCIT, we propose a dynamic knowledge organization to effectively integrate updates from different tasks during training and subspace selective activation to allocate task-specific output during inference. Extensive experimental results demonstrate that our proposed method significantly enhances model performance across varying levels of data heterogeneity and catastrophic forgetting. Code and dataset are released at https://github.com/Ghy0501/FCIT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5bae733f-7017-47f1-ada0-ddb830865216Cited by top-tier papers5
- RobustMerge: Parameter-Efficient Model Merging for MLLMs with Direction RobustnessFanhu Zeng, Haiyang Guo, Fei Zhu, Li Shen et al.NeurIPS 2025 · 28 citations
- Multimodal Continual Instruction Tuning with Dynamic Gradient GuidanceSongze Li, Mingyu Gao, Tonghua Su, Xu-Yao Zhang et al.CVPR 2026 · 6 citations
- φ-DPO: Fairness Direct Preference Optimization Approach to Continual Learning in Large Multimodal ModelsThanh-Dat Truong, Huu-Thien Tran, Jackson David Cothren, Bhiksha Raj et al.CVPR 2026 · 2 citations
- Model-Dowser: Data-Free Importance Probing to Mitigate Catastrophic Forgetting in Multimodal Large Language ModelsHyeontaek Hwang, DINH SON NGUYEN, Daeyoung KimICML 2026
- One Patch Doesn't Fit All: Adaptive Patching for Native-Resolution Multimodal Large Language ModelsWenzhuo Liu, Weijie Yin, Fei Zhu, Shijie Ma et al.ICLR 2026
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
Related papers
- SMoLoRa: Exploring and Defying Dual Catastrophic Forgetting in Continual Visual Instruction TuningZiqi Wang, Chang Che, Qi Wang, Yangyang Li et al.ICCV 2025 · 4 citations
- Advancing SMoE for Continuous Domain Adaptation of MLLMs: Adaptive Router and Domain-Specific LossLiang Zhang, Ziyao Lu, Fandong Meng, Hui Li et al.ACL 2025 · 3 citations
- ModalPrompt: Towards Efficient Multimodal Continual Instruction Tuning with Dual-Modality Guided PromptFanhu Zeng, Fei Zhu, Haiyang Guo, Xu-Yao Zhang et al.EMNLP 2025
- HiDe-LLaVA: Hierarchical Decoupling for Continual Instruction Tuning of Multimodal Large Language ModelHaiyang Guo, Fanhu Zeng, Ziwei Xiang, Fei Zhu et al.ACL 2025
- Don't Half-listen: Capturing Key-part Information in Continual Instruction TuningYongquan He, Wenyuan Zhang, Xuancheng Huang, Peng Zhang et al.ACL 2025
