Where and What Matters: Sensitivity-Aware Task Vectors for Many-Shot Multimodal In-Context Learning
Ziyu Ma, Chenhui Gou, Yiming Hu, Yong Wang, Bohan Zhuang, Jianfei Cai
Abstract
Large Multimodal Models (LMMs) have shown promising in-context learning (ICL) capabilities, but scaling to manyshot settings remains difficult due to limited context length and high inference cost. To address these challenges, taskvector-based methods have been explored by inserting compact representations of many-shot in-context demonstrations into model activations. However, existing task-vector-based methods either overlook the importance of where to insert task vectors or struggle to determine suitable values for each location. To this end, we propose a novel Sensitivityaware Task Vector insertion framework (STV) to figure out where and what to insert. Our key insight is that activation deltas across query-context pairs exhibit consistent structural patterns, providing a reliable cue for insertion. Based on the identified sensitive-aware locations, we construct a pre-clustered activation bank for each location by clustering the activation values, and then apply reinforcement learning to choose the most suitable one to insert. We evaluate STV across a range of multimodal models (e.g., Qwen-VL, Idefics-2) and tasks (e.g., VizWiz, OK-VQA), demonstrating its effectiveness and showing consistent improvements over previous task-vector-based methods with strong generalization. Our code will be available at https://github.com/AMAP- ML/STV.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f6d80301-421a-4037-bb09-9c3f4b226f77Cited by top-tier papers2
- CoEvolve: Training LLM Agents via Agent-Data Mutual EvolutionShidong Yang, Ziyu Ma, Tongwen Huang, Yiming Hu et al.ACL 2026 · 6 citations
- VQ-VA World: Towards High-Quality Visual Question-Visual AnsweringChenhui Gou, Zilong Chen, Zeyu Wang, Feng Li et al.CVPR 2026 · 3 citations
Builds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Implicit In-context LearningZhuowei Li, Zihao Xu, Ligong Han, Yunhe Gao et al.ICLR 2025 · 1,989 citations
- YaRN: Efficient Context Window Extension of Large Language ModelsBowen Peng, Jeffrey Quesnelle, Honglu Fan, Enrico ShippoleICLR 2024 · 508 citations
- What matters when building vision-language models?Hugo Laurençon, Léo Tronchon, Matthieu Cord, Victor SanhNeurIPS 2024 · 401 citations
Related papers
- Mimic In-Context Learning for Multimodal TasksYuchu Jiang, Jiale Fu, Chenduo Hao, Xinting Hu et al.CVPR 2025
- Multimodal Task Vectors Enable Many-Shot Multimodal In-Context LearningBrandon Huang, Chancharik Mitra, Leonid Karlinsky, Assaf Arbelle et al.NeurIPS 2024 · 60 citations
- LIVE: Learnable In-Context Vector for Visual Question AnsweringYingzhe Peng, Chenduo Hao, Xinting Hu, Jiawei Peng et al.NeurIPS 2024
- Provoking Multi-modal Few-Shot LVLM via Exploration-Exploitation In-Context LearningCheng Chen, Yunpeng Zhai, Yifan Zhao, Jinyang Gao et al.CVPR 2025
- Beyond Demonstrations: Dynamic Vector Construction from Latent RepresentationsWang Cai, Hsiu-Yuan Huang, Zhixiang Wang, Yunfang WuEMNLP 2025
