ECA: Efficient Continual Alignment for Open-Ended Image-to-Text Generation
Jiangtao Kong, Peijun Zhao, Chun-Fu (Richard) Chen, Youngwook Do, Shaohan Hu, Tianyi Zhou, Huajie Shao
Abstract
Incremental Learning (IL) for Open-ended Image-to-Text Generation (OpenITG) enables models to continuously generate accurate, contextually relevant text for new images while preserving previously acquired knowledge. Unlike prior studies, this paper addresses a more practical scenario in which the predominant category of visual data shifts over time as environments evolve. In this context, we introduce a new notion of continual alignment, which incrementally adapts the alignment module within pre-trained VLMs to preserve high-quality cross-modal representations. Based on this idea, we propose E fficient C ontinual A lignment (ECA), a novel exemplar-free IL approach for OpenITG. The key challenge is enabling the model to acquire new, task-specific features while minimizing interference with the established alignment without accessing raw data from previous tasks. To address this, ECA employs three core mechanisms: a M ixture o f Q uery (MoQ) module that adapts task-specific query tokens, a F ish e r D ynamic Ex pansion (FeDEx) that dynamically expands model structure based on a Fisher Information Matrix (FIM)-based metric, and an embedding dictionary with D ictionary R eplay (DR) to retain past knowledge. To evaluate ECA's performance, we construct four new IL OpenITG benchmarks that better reflect real-world scenarios. Experimental results demonstrate that ECA significantly mitigates catastrophic forgetting and improves IL performance compared to baseline methods. Code and benchmarks are available at https://github.com/Snowball0823/ECA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d4aaeb6b-d609-462b-8378-2f9173996c0cBuilds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- Breaking the Synthetic-Real Domain Shortcut for Training-Free Generative Replay-based Class Incremental LearningTao Zhang, Qixuan Fan, Yiyuan Liang, Yanjie Wang et al.ICML 2026
- Embracing Language Inclusivity and Diversity in CLIP through Continual Language LearningBang Yang, Yong Dai, Xuxin Cheng, Yaowei Li et al.AAAI 2024 · 9 citations
- Synthetic Data is an Elegant GIFT for Continual Vision-Language ModelsBin Wu, Wuxuan Shi, Jinqiao Wang, Mang YeCVPR 2025
- Dynamic Multi-Layer Null Space Projection for Vision-Language Continual LearningBorui Kang, Lei Wang, Zhiping Wu, Tao Feng et al.ICCV 2025 · 4 citations
- Pi-CCA: Prompt-Invariant CCA Certificates for Replay-Free Continual Multimodal LearningJiayu Zhang, Chuangxin Zhao, Canran Xiao, Ruibo Duan et al.ICLR 2026
