How to Configure Good In-Context Sequence for Visual Question Answering
Li Li, Jiawei Peng, Huiyi Chen, Chongyang Gao, Xu Yang
摘要
Inspired by the success of Large Language Models in dealing with new tasks via In-Context Learning (ICL) in NLP, researchers have also developed Large Vision-Language Models (LVLMs) with ICL capabilities. However, when implementing ICL using these LVLMs, researchers usually resort to the simplest way like random sampling to configure the in-context sequence, thus leading to suboptimal results. To enhance the ICL performance, in this study, we use Visual Question Answering (VQA) as case study to explore diverse in-context configurations to find the powerful ones. Additionally, through observing the changes of the LVLM outputs by altering the in-context sequence, we gain insights into the inner properties of LVLMs, improving our understanding of them. Specifically, to explore incontext configurations, we design diverse retrieval methods and employ different strategies to manipulate the retrieved demonstrations. Through exhaustive experiments on three VQA datasets: VQAv2, VizWiz, and OK-VQA, we uncover three important inner properties of the applied LVLM and demonstrate which strategies can consistently improve the ICL VQA performance. Our code is provided in: https: //github.com/GaryJiajia/OFv2_ICL_VQA .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper31
- Bridge the Modality and Capability Gaps in Vision-Language Model SelectionChao Yi, Yuhang He, De-Chuan Zhan, Han-Jia YeNeurIPS 2024 · 被引用 32 次
- MOS: Model Surgery for Pre-Trained Model-Based Class-Incremental LearningHai-Long Sun, Da-Wei Zhou, Hanbin Zhao, Le Gan 等AAAI 2025 · 被引用 31 次
- MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMsHuiyi Chen, Jiawei Peng, Dehai Min, Changchang Sun 等ICML 2026 · 被引用 18 次
- TD3: Tucker Decomposition Based Dataset Distillation Method for Sequential RecommendationJiaqing Zhang, Mingjia Yin, Hao Wang, Yawen Li 等WWW 2025 · 被引用 17 次
- Lever LM: Configuring In-Context Sequence to Lever Large Vision Language ModelsXu Yang, Yingzhe Peng, Haoxuan Ma, Shuo Xu 等NeurIPS 2024 · 被引用 13 次
它引用的顶会 Paper19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
相关 Paper
- Provoking Multi-modal Few-Shot LVLM via Exploration-Exploitation In-Context LearningCheng Chen, Yunpeng Zhai, Yifan Zhao, Jinyang Gao 等CVPR 2025
- What Factors Affect Multi-Modal In-Context Learning? An In-Depth ExplorationLibo Qin, Qiguang Chen, Hao Fei, Zhi Chen 等NeurIPS 2024 · 被引用 37 次
- In-Context Compositional Generalization for Large Vision-Language ModelsChuanhao Li, Chenchen Jing, Zhen Li, Mingliang Zhai 等EMNLP 2024 · 被引用 1 次
- Exploring Diverse In-Context Configurations for Image CaptioningXu Yang, Yongliang Wu, Mingzhuo Yang, Haokun Chen 等NeurIPS 2023 · 被引用 104 次
- CCL: Causal-aware In-context Learning for Out-of-Distribution GeneralizationHoyoon Byun, Gyeongdeok Seo, Joonseong Kang, Taero Kim 等NeurIPS 2025 · 被引用 1 次
