Adapting Multi-modal Large Language Model to Concept Drift From Pre-training Onwards
Xiaoyu Yang, Jie Lu, En Yu
摘要
Multi-modal Large Language Models (MLLMs) frequently face challenges from concept drift when dealing with real-world streaming data, wherein distributions change unpredictably. This mainly includes gradual drift due to long-tailed data and sudden drift from Out-Of-Distribution (OOD) data, both of which have increasingly drawn the attention of the research community. While these issues have been extensively studied in the individual domain of vision or language, their impacts on MLLMs in concept drift settings remain largely underexplored. In this paper, we reveal the susceptibility and vulnerability of Vision-Language (VL) models to significant biases arising from gradual drift and sudden drift, particularly in the pre-training. To effectively address these challenges, we propose a unified framework that extends concept drift theory to the multi-modal domain, enhancing the adaptability of the VL model to unpredictable distribution changes. Additionally, a T-distribution based drift adapter is proposed to effectively mitigate the bias induced by the gradual drift, which also facilitates the model in distinguishing sudden distribution changes through explicit distribution modeling. Extensive experiments demonstrate our method enhances the efficiency and accuracy of image-text alignment in the pre-training of VL models, particularly in the concept drift scenario. Moreover, various downstream tasks exhibit significant improvements in our model's ability to adapt to the long-tailed open world. Furthermore, we create a set of multi-modal datasets called OpenMMlo, specifically tailored for the long-tailed open-world setting, to validate our findings. To foster the development of the multi-modal community, we have made both OpenMMlo datasets and our code publicly available at: https://github.com/XiaoyuYoung/ConceptDriftMLLMs .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Walking the Tightrope: Autonomous Disentangling Beneficial and Detrimental Drifts in Non-Stationary Custom-TuningXiaoyu Yang, Jie Lu, En YuNeurIPS 2025 · 被引用 22 次
- Learning Robust Spectral Dynamics for Temporal Domain GeneralizationEn Yu, Jie Lu, Xiaoyu Yang, Guangquan Zhang 等NeurIPS 2025 · 被引用 22 次
- Drift-aware Collaborative Assistance Mixture of Experts for Heterogeneous Multistream LearningEn Yu, Jie Lu, Kun Wang, Xiaoyu Yang 等AAAI 2026 · 被引用 15 次
- From Newborn to Impact: Bias-Aware Citation PredictionMingfei Lu, Mengjia Wu, Jiawei Xu, Weikai Li 等WWW 2026 · 被引用 6 次
- Generalized Incremental Learning under Concept Drift across Evolving Data StreamsEn Yu, Jie Lu, Guangquan ZhangWWW 2026 · 被引用 4 次
它引用的顶会 Paper41
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
相关 Paper
- Lyapunov-Stable Adaptive Control for Multimodal Concept DriftTianyu Bell Pan, Mengdi Zhu, Alexa Jordyn Cole, Ronald Wilson 等NeurIPS 2025
- Δ Energy: Optimizing Energy Change During Vision-Language Alignment Improves both OOD Detection and OOD GeneralizationLin Zhu, Yifeng Yang, Xinbing Wang, Qinying Gu 等NeurIPS 2025 · 被引用 2 次
- Online Drift Detection with Maximum Concept DiscrepancyKe Wan, Yi Liang, Susik YoonKDD 2024 · 被引用 4 次
- Test-Time Retrieval-Augmented Adaptation for Vision-Language ModelsXinqi Fan, Xueli Chen, Luoxiao Yang, Chuin Hong Yap 等ICCV 2025 · 被引用 4 次
- Dual Prototype Evolving for Test-Time Generalization of Vision-Language ModelsCe Zhang, Simon Stepputtis, Katia P. Sycara, Yaqi XieNeurIPS 2024 · 被引用 57 次
