Why is a Bird's Caption a Good Demonstration? Towards Effective Multimodal In-Context Learning without Dedicated Data
Junlin Fang, Wenya Wang, Lingli Zhang, Fengmao Lv
Abstract
Multimodal Large Language Models (MLLMs) have achieved impressive performance across a range of tasks by leveraging Multimodal In-Context Learning (MICL), which uses a few task-specific examples as demonstrations. However, existing approaches assume the availability of pre-prepared curated datasets that serve as support sets, limiting the adaptability of MICL to novel and unseen tasks where dedicated data is unavailable. To fill this research gap, we first explore the effectiveness of MICL using non-customized data. Through systematic evaluations across 17 datasets and five state-of-the-art MLLMs, we demonstrate significant performance gains with MICL compared to zero-shot evaluation. To more thoroughly understand underlying reasons behind this phenomenon, we posit and validate two hypotheses: 1) multimodal demonstrations facilitate cross-modal interactions and 2) demonstrations provide transferable knowledge. Building on these insights, we explore factors that affect MICL and arrive at several key takeaways. First, to address the limitations of existing retrieval methods in MICL without dedicated data, we propose a Fast Maximum Mean Discrepancy based (FMMD) retrieval metric and a Semantics-Modality Relation-Aware (SMRA) retrieval metric to perform inter- and intra-dataset retrieval, respectively. Additionally, we find that increasing demonstrations, combining demonstrations from diverse datasets, and providing instructions for query samples can further boost MICL. We hope this study can inspire future works on improving MICL in real-world scenarios.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 018d5ad5-ea8e-466a-8f0a-243f3bf4a0e1Related papers
- An Empirical Study on Configuring In-Context Learning Demonstrations for Unleashing MLLMs' Sentimental Perception CapabilityDaiqing Wu, Dongbao Yang, Sicheng Zhao, Can Ma et al.ICML 2025
- What Factors Affect Multi-Modal In-Context Learning? An In-Depth ExplorationLibo Qin, Qiguang Chen, Hao Fei, Zhi Chen et al.NeurIPS 2024 · 37 citations
- Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and BottlenecksYu Wang, Sharon LiACL 2026
- Link-Context Learning for Multimodal LLMsYan Tai, Weichen Fan, Zhao Zhang, Ziwei LiuCVPR 2024 · 7 citations
- Mixture of Demonstrations for In-Context LearningSong Wang, Zihan Chen, Chengshuai Shi, Cong Shen et al.NeurIPS 2024 · 21 citations
