Branch, or Layer? Zeroth-Order Optimization for Continual Learning of Vision-Language Models
Ziwei Liu, Borui Kang, Wei Li, Hangjie Yuan, Yanbing Yang, Wenbin Li, Yifan Zhu, Tao Feng, Jun Luo
Abstract
Vision-Language Continual Learning (VLCL) has attracted significant research attention for its robust capabilities, and the adoption of Parameter-Efficient Fine-Tuning (PEFT) strategies is enabling these models to achieve competitive performance with substantially reduced resource consumption. However, dominated First-Order (FO) optimization is prone to trap models in suboptimal local minima, especially in limited exploration subspace within PEFT. To overcome this challenge, this paper pioneers a systematic exploration of adopting Zeroth-Order (ZO) optimization for PEFT-based VLCL. We first identify the incompatibility of naive full-ZO adoption in VLCL due to optimization process instability. We then investigate the application of ZO optimization from a modality branch-wise to a fine-grained layer-wise across various training units to identify an optimal strategy. Besides, a key theoretical insight reveals that vision modality exhibit higher variance than language counterparts in VLCL during the ZO optimization process, and we propose a modality-aware stabilized ZO strategy, which adopts gradient sign normalization in ZO and constrains vision modality perturbation to further improve performance. Benefiting from the adoption of ZO optimization, PEFT-based VLCL fulfills better ability to escape local minima during the optimization process, extensive experiments on four benchmarks demonstrate that our method achieves state-of-the-art results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Rep Deep & Machine Learning: Exemplar-Free Continual Video Action Recognition via Slow-Fast Collaborative LearningXueyi Zhang, Chengwei Zhang, Zheng Li, Xiyu Wang et al.AAAI 2026 · 1 citation
- Don't Forget Why You Started: Tackling Dual Forgetting in Vision-Language Continual LearningBorui Kang, Jinrui Gu, Tao Feng, Qi Fan et al.ICML 2026
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation LearningWeixin Liang, Yuhui Zhang, Yongchan Kwon, Serena Yeung et al.NeurIPS 2022 · 834 citations
- Learning to Prompt for Continual LearningZifeng Wang, Zizhao Zhang, Chen-Yu Lee, Han Zhang et al.CVPR 2022 · 635 citations
- Fine-Tuning Language Models with Just Forward PassesSadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian et al.NeurIPS 2023 · 495 citations
- Gradient Projection Memory for Continual LearningGobinda Saha, Isha Garg, Kaushik RoyICLR 2021 · 409 citations
Related papers
- Bilevel ZOFO: Efficient LLM Fine-Tuning and Meta-TrainingReza Shirkavand, Peiran Yu, Qi He, Heng HuangNeurIPS 2025 · 6 citations
- Fed-Duet: Dual Expert-Orchestrated Framework for Continual Federated Vision-Language LearningTao Guo, Junwei Chen, Laizhong CuiICLR 2026
- Enhancing Zeroth-order Fine-tuning for Language Models with Low-rank StructuresYiming Chen, Yuan Zhang, Liyuan Cao, Kun Yuan et al.ICLR 2025
- New Hybrid Fine-Tuning Paradigm for LLMs: Algorithm Design and Convergence Analysis FrameworkShaocong Ma, Peiran Yu, Heng HuangICLR 2026 · 1 citation
- Boosting Continual Learning of Vision-Language Models via Mixture-of-Experts AdaptersJiazuo Yu, Yunzhi Zhuge, Lu Zhang, Ping Hu et al.CVPR 2024 · 80 citations
