Test-Time Visual In-Context Tuning
Jiahao Xie, Alessio Tonioni, Nathalie Rauschmayr, Federico Tombari, Bernt Schiele
Abstract
Figure 1 . Test-time visual in-context tuning (VICT) on six representative vision tasks under distribution shifts. We benchmark the robustness of VICL with 15 common corruptions adopted in [23, 31] , and report the averaged performance across all corruptions. Existing VICL models like Painter exhibit poor generalization capability to unseen new domains when the task prompts come from the training distribution (i.e., zero-shot). Performances are even worse when given task prompts from the test distribution (i.e., one-shot). By performing VICT at test time, we can significantly improve Painter in both zero-shot and one-shot manners.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 61fa09f1-9329-4c15-9517-7ecc118aec67Cited by top-tier papers1
Ask how each one uses itBuilds on24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
Related papers
- VPA: Fully Test-Time Visual Prompt AdaptationJiachen Sun, Mark Ibrahim, Melissa Hall, Ivan Evtimov et al.ACM MM 2023 · 7 citations
- Images Speak in Images: A Generalist Painter for In-Context Visual LearningXinlong Wang, Wen Wang, Yue Cao, Chunhua Shen et al.CVPR 2023
- Test-Time Prompt Tuning for Zero-Shot Generalization in Vision-Language ModelsManli Shu, Weili Nie, De-An Huang, Zhiding Yu et al.NeurIPS 2022 · 603 citations
- Align Your Prompts: Test-Time Prompting with Distribution Alignment for Zero-Shot GeneralizationJameel Abdul Samadh, Hanan Gani, Noor Hussein, Muhammad Uzair Khattak et al.NeurIPS 2023 · 147 citations
- Patching open-vocabulary models by interpolating weightsGabriel Ilharco, Mitchell Wortsman, Samir Yitzhak Gadre, Shuran Song et al.NeurIPS 2022 · 230 citations
