Provoking Multi-modal Few-Shot LVLM via Exploration-Exploitation In-Context Learning
Cheng Chen, Yunpeng Zhai, Yifan Zhao, Jinyang Gao, Bolin Ding, Jia Li
Abstract
In-context learning (ICL), a predominant trend in instruction learning, aims at enhancing the performance of large language models by providing clear task guidance and examples, improving their capability in task understanding and execution. This paper investigates ICL on Large Vision-Language Models (LVLMs) and explores the policies of multi-modal demonstration selection. Existing research efforts in ICL face significant challenges: First, they rely on pre-defined demonstrations or heuristic selecting strategies based on human intuition, which are usually inadequate for covering diverse task requirements, leading to sub-optimal solutions; Second, individually selecting each demonstration fails in modeling the interactions between them, resulting in information redundancy. Unlike these prevailing efforts, we propose a new exploration-exploitation reinforcement learning framework, which explores policies to fuse multi-modal information and adaptively select adequate demonstrations as an integrated whole. The framework allows LVLMs to optimize themselves by continually refining their demonstrations through self-exploration, enabling the ability to autonomously identify and generate the most effective selection policies for in-context learning. Experimental results verify the superior performance of our approach on four Visual Question-Answering (VQA) datasets, demonstrating its effectiveness in enhancing the generalization capability of few-shot LVLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5349463b-d5aa-4ad5-ba16-0f4fd5b2f0c4Cited by top-tier papers2
- ContextNav: Towards Agentic Multimodal In-Context LearningHonghao Fu, Yuan Ouyang, Kai-Wei Chang, Yiwei Wang et al.ICLR 2026 · 14 citations
- Retrieving Counterfactuals Improves Visual In-Context LearningGuangzhi Xiong, Sanchit Sinha, Zhenghao He, Aidong ZhangCVPR 2026 · 3 citations
Builds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order SensitivityYao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel et al.ACL 2022 · 1,494 citations
Related papers
- How to Configure Good In-Context Sequence for Visual Question AnsweringLi Li, Jiawei Peng, Huiyi Chen, Chongyang Gao et al.CVPR 2024
- What Factors Affect Multi-Modal In-Context Learning? An In-Depth ExplorationLibo Qin, Qiguang Chen, Hao Fei, Zhi Chen et al.NeurIPS 2024 · 37 citations
- In-Context Compositional Generalization for Large Vision-Language ModelsChuanhao Li, Chenchen Jing, Zhen Li, Mingliang Zhai et al.EMNLP 2024 · 1 citation
- Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and BottlenecksYu Wang, Sharon LiACL 2026
- Active Example Selection for In-Context LearningYiming Zhang, Shi Feng, Chenhao TanEMNLP 2022 · 84 citations
