Make LVLMs Focus: Context-Aware Attention Modulation for Better Multimodal In-Context Learning
Yanshu Li, Jianjiang Yang, Ziteng Yang, Bozheng Li, Ligong Han, Hongyang He, Zhengtao Yao, Yingjie Victor Chen, Songlin Fei, Dongfang Liu, Ruixiang Tang
Abstract
Multimodal in-context learning (ICL) is becoming a key capability that allows large vision-language models (LVLMs) to adapt to novel tasks without parameter updates, which expands their usefulness in many real-world applications. However, ICL performance remains unstable even when the in-context demonstrations (ICDs) are well matched, showing that LVLMs still struggle to make full use of the provided context. While existing work mainly focuses on prompt engineering or post-hoc logit calibration, we study the attention mechanisms inside LVLMs to address their inherent limitations. We identify two important weaknesses in their self-attention that hinder effective ICL. To address these weaknesses, we propose Context-Aware Modulated Attention (CAMA), a training-free and plug-and-play method that dynamically adjusts attention logits based on the input in-context sequence. CAMA uses a two-stage modulation process that strengthens attention to semantically important tokens, especially visual ones. Across four LVLMs and seven benchmarks, CAMA consistently outperforms vanilla models and baselines, showing clear effectiveness and generalization. It can also activate the intended benefits of prompt engineering methods and remains robust across different sequence configurations. Therefore, CAMA opens up new directions for improving multimodal reasoning through a deeper understanding of attention dynamics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3b80cffd-3931-41b0-93e8-c72a85236186Cited by top-tier papers10
- CADMorph: Geometry‑Driven Parametric CAD Editing via a Plan-Generate-Verify LoopWeijian Ma, Shizhao Sun, Ruiyu Wang, Jiang BianNeurIPS 2025 · 4 citations
- DEALT: LLM-driven Diversity-Enhanced Data Augmentation for Long-Tail Text ClassificationWayne Lu, Xiaoxi CuiAAAI 2026 · 2 citations
- Semantic Granularity Navigation in Image EditingLiangsi Lu, Minzhe Guo, Xuhang Chen, Yang ShiICML 2026 · 1 citation
- Point Cloud Self-Supervised Learning via 3D to Multi-View Masked LearnerZhimin Chen, Xuewei Chen, Xiao Guo, Yingwei Li et al.ICCV 2025 · 1 citation
- HiFICL: High-Fidelity In-Context Learning for Multimodal TasksXiaoyu Li, Yuhang Liu, xuanshuo kang, zheng luo et al.CVPR 2026 · 1 citation
Builds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- The Learnability of In-Context LearningNoam Wies, Yoav Levine, Amnon ShashuaNeurIPS 2023 · 207 citations
- Exploring Diverse In-Context Configurations for Image CaptioningXu Yang, Yongliang Wu, Mingzhuo Yang, Haokun Chen et al.NeurIPS 2023 · 104 citations
Related papers
- TACO: Enhancing Multimodal In-context Learning via Task Mapping-Guided Sequence ConfigurationYanshu Li, Jianjiang Yang, Tian Yun, Pinyuan Feng et al.EMNLP 2025 · 2 citations
- Enhancing Retrieval-Augmented Large Vision Language Models via Knowledge Conflict MitigationWenbin An, Jiahao Nie, Feng Tian, Mingxiang Cai et al.AAAI 2026
- Interleaved-Modal Chain-of-ThoughtJun Gao, Yongqi Li, Ziqiang Cao, Wenjie LiCVPR 2025
- Mitigating Object Hallucinations in Large Vision-Language Models with Assembly of Global and Local AttentionWenbin An, Feng Tian, Sicong Leng, Jiahao Nie et al.CVPR 2025
- Hyper-ICL: Attention Calibration with Hyperbolic Anchor Distillation for Multimodal In-Context LearningNiloufar Alipour Talemi, Hossein Kashiani, Fatemeh AfghahICML 2026
