Why is a Bird's Caption a Good Demonstration? Towards Effective Multimodal In-Context Learning without Dedicated Data
Junlin Fang, Wenya Wang, Lingli Zhang, Fengmao Lv
摘要
Multimodal Large Language Models (MLLMs) have achieved impressive performance across a range of tasks by leveraging Multimodal In-Context Learning (MICL), which uses a few task-specific examples as demonstrations. However, existing approaches assume the availability of pre-prepared curated datasets that serve as support sets, limiting the adaptability of MICL to novel and unseen tasks where dedicated data is unavailable. To fill this research gap, we first explore the effectiveness of MICL using non-customized data. Through systematic evaluations across 17 datasets and five state-of-the-art MLLMs, we demonstrate significant performance gains with MICL compared to zero-shot evaluation. To more thoroughly understand underlying reasons behind this phenomenon, we posit and validate two hypotheses: 1) multimodal demonstrations facilitate cross-modal interactions and 2) demonstrations provide transferable knowledge. Building on these insights, we explore factors that affect MICL and arrive at several key takeaways. First, to address the limitations of existing retrieval methods in MICL without dedicated data, we propose a Fast Maximum Mean Discrepancy based (FMMD) retrieval metric and a Semantics-Modality Relation-Aware (SMRA) retrieval metric to perform inter- and intra-dataset retrieval, respectively. Additionally, we find that increasing demonstrations, combining demonstrations from diverse datasets, and providing instructions for query samples can further boost MICL. We hope this study can inspire future works on improving MICL in real-world scenarios.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- An Empirical Study on Configuring In-Context Learning Demonstrations for Unleashing MLLMs' Sentimental Perception CapabilityDaiqing Wu, Dongbao Yang, Sicheng Zhao, Can Ma 等ICML 2025
- What Factors Affect Multi-Modal In-Context Learning? An In-Depth ExplorationLibo Qin, Qiguang Chen, Hao Fei, Zhi Chen 等NeurIPS 2024 · 被引用 37 次
- Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and BottlenecksYu Wang, Sharon LiACL 2026
- Link-Context Learning for Multimodal LLMsYan Tai, Weichen Fan, Zhao Zhang, Ziwei LiuCVPR 2024 · 被引用 7 次
- Mixture of Demonstrations for In-Context LearningSong Wang, Zihan Chen, Chengshuai Shi, Cong Shen 等NeurIPS 2024 · 被引用 21 次
