AIM: Let Any Multimodal Large Language Models Embrace Efficient In-Context Learning
Jun Gao, Qian Qiao, Tianxiang Wu, Zili Wang, Ziqiang Cao, Wenjie Li
Abstract
In-context learning (ICL) advances Large Language Models (LLMs) exhibiting emergent ability on downstream tasks without updating billions of parameters. However, in the area of multimodal Large Language Models (MLLMs), two problems hinder the application of multimodal ICL: (1) Most primary MLLMs are only trained on single-image datasets, making them unable to read extra multimodal demonstrations. (2) With the demonstrations increasing, thousands of visual tokens highly challenge hardware and degrade ICL performance. During preliminary explorations, we discovered that the inner LLM focuses more on the linguistic modality within multimodal demonstrations during generation. Therefore, we propose a general and light-weighted framework AIM to tackle the mentioned problems through Aggregating Image information of Multimodal demonstrations to the latent space of the corresponding textual labels. After aggregation, AIM substitutes each demonstration with generated fused virtual tokens whose length is reduced to the same as its texts. Except for shortening input length, AIM further upgrades MLLMs pre-trained on image-text pairs to support multimodal ICL, as images from demonstrations are disregarded. Furthermore, benefiting from aggregating different demonstrations independently, AIM configures Demonstration Bank (DB) to avoid repeated aggregation, which significantly boosts model efficiency. We build AIM upon QWen-VL and LLaVA-Next, and AIM is comprehensively evaluated on image caption, VQA, and hateful speech detection. Outstanding results reveal that AIM provides an efficient and effective solution in upgrading MLLMs for multimodal ICL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6edb490e-4a83-4402-b2f6-a836a56f34d9Cited by top-tier papers5
- See the Forest for the Trees: Loosely Speculative Decoding via Visual-Semantic Guidance for Efficient Inference of Video LLMsYicheng Ji, Jun Zhang, Jinpeng Chen, Cong Wang et al.ACL 2026 · 4 citations
- Towards Predicting Any Human Trajectory In ContextRyo Fujii, Hideo Saito, Ryo HachiumaNeurIPS 2025 · 2 citations
- VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video UnderstandingKangsan Kim, Geon Park, Youngwan Lee, Woongyeong Yeo et al.CVPR 2025
- Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and BottlenecksYu Wang, Sharon LiACL 2026
- PointLLM-R: Enhancing 3D Point Cloud Reasoning via Chain-of-ThoughtChaoqi Chen, Qile Xu, Wenjun Zhou, Hui HuangSIGGRAPH 2026
Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
Related papers
- Provoking Multi-modal Few-Shot LVLM via Exploration-Exploitation In-Context LearningCheng Chen, Yunpeng Zhai, Yifan Zhao, Jinyang Gao et al.CVPR 2025
- In-Context Compositional Generalization for Large Vision-Language ModelsChuanhao Li, Chenchen Jing, Zhen Li, Mingliang Zhai et al.EMNLP 2024 · 1 citation
- LIVE: Learnable In-Context Vector for Visual Question AnsweringYingzhe Peng, Chenduo Hao, Xinting Hu, Jiawei Peng et al.NeurIPS 2024
- MMICL: Empowering Vision-language Model with Multi-Modal In-Context LearningHaozhe Zhao, Zefan Cai, Shuzheng Si, Xiaojian Ma et al.ICLR 2024 · 206 citations
- Improved Baselines with Visual Instruction TuningHaotian Liu, Chunyuan Li, Yuheng Li, Yong Jae LeeCVPR 2024
