PromptHub: Enhancing Multi-Prompt Visual In-Context Learning with Locality-Aware Fusion, Concentration and Alignment
Tianci Luo, Jinpeng Wang, Shiyu Qin, Niu Lian, Yan Feng, Bin Chen, Chun Yuan, Shu-Tao Xia
Abstract
Visual In-Context Learning (VICL) aims to complete vision tasks by imitating pixel demonstrations. Recent work (Wang et al., 2025) pioneered prompt fusion that combines the advantages of various demonstrations, which shows a promising way to extend VICL. Unfortunately, the patch-wise fusion framework and model-agnostic supervision hinder the exploitation of informative cues, thereby limiting performance gains. To overcome this deficiency, we introduce PromptHub, a framework that holistically strengthens multi-prompting through locality-aware fusion, concentration and alignment. PromptHub exploits spatial priors to capture richer contextual information, employs complementary concentration, alignment, and prediction objectives to mutually guide training, and incorporates data augmentation to further reinforce supervision. Extensive experiments on three fundamental vision tasks demonstrate the superiority of PromptHub. Moreover, we validate its universality, transferability, and robustness across out-of-distribution settings, and various retrieval scenarios. This work establishes a reliable locality-aware paradigm for prompt fusion, moving beyond prior patch-wise approaches. Code is available at https://github.com/luotc-why/ICLR26-PromptHub .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ffe67344-b975-4751-9674-beee674c4d01Cited by top-tier papers2
- Learning Cross-View Object Correspondence via Cycle-Consistent Mask PredictionShannan Yan, Leqi Zheng, Keyu Lv, Jingchen Ni et al.CVPR 2026 · 5 citations
- Love Me, Love My Label: Rethinking the Role of Labels in Prompt Retrieval for Visual In-Context LearningTianci Luo, Haohao Pan, Jinpeng Wang, Niu Lian et al.CVPR 2026
Builds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento et al.ICML 2023 · 729 citations
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng et al.EMNLP 2024 · 479 citations
- Visual Prompting via Image InpaintingAmir Bar, Yossi Gandelsman, Trevor Darrell, Amir Globerson et al.NeurIPS 2022 · 340 citations
Related papers
- CoPL: Contextual Prompt Learning for Vision-Language UnderstandingKoustava Goswami, Srikrishna Karanam, Prateksha Udhayanan, K. J. Joseph et al.AAAI 2024 · 20 citations
- Towards Global Optimal Visual In-Context Learning Prompt SelectionChengming Xu, Chen Liu, Yikai Wang, Yuan Yao et al.NeurIPS 2024 · 19 citations
- In-context Prompt-augmented Micro-video Popularity PredictionZhangtao Cheng, Jiao Li, Jian Lang, Ting Zhong et al.AAAI 2025 · 3 citations
- Enhancing Pre-trained ViTs for Downstream Task Adaptation: A Locality-Aware Prompt Learning MethodShaokun Wang, Yifan Yu, Yuhang He, Yihong GongACM MM 2024 · 1 citation
- FOCoOp: Enhancing Out-of-Distribution Robustness in Federated Prompt Learning for Vision-Language ModelsXinting Liao, Weiming Liu, Jiaming Qian, Pengyang Zhou et al.ICML 2025
