SVIP: Semantically Contextualized Visual Patches for Zero-Shot Learning
Zhi Chen, Zecheng Zhao, Jingcai Guo, Jingjing Li, Zi Huang
Abstract
Zero-shot learning (ZSL) aims to recognize unseen classes without labeled training examples by leveraging class-level semantic descriptors such as attributes. A fundamental challenge in ZSL is semantic misalignment, where semantic-unrelated information involved in visual features introduce ambiguity to visual-semantic interaction. Unlike existing methods that suppress semantic-unrelated information post hoc either in the feature space or the model space, we propose addressing this issue at the input stage, preventing semantic-unrelated patches from propagating through the network. To this end, we introduce Semantically contextualized VIsual Patches (SVIP) for ZSL, a transformer-based framework designed to enhance visual-semantic alignment. Specifically, we propose a self-supervised patch selection mechanism that preemptively learns to identify semantic-unrelated patches in the input space. This is trained with the supervision from aggregated attention scores across all transformer layers, which estimate each patch's semantic score. As removing semantic-unrelated patches from the input sequence may disrupt object structure, we replace them with learnable patch embeddings. With initialization from word embeddings, we can ensure they remain semantically meaningful throughout feature extraction. Extensive experiments on ZSL benchmarks demonstrate that SVIP achieves state-of-the-art performance results while providing more interpretable and semantically rich feature representations. Code is available at https://github.com/uqzhichen/SVIP.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5fb469db-2767-426b-ba44-3916f52a07eeCited by top-tier papers1
Ask how each one uses itBuilds on33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetLi Yuan, Yunpeng Chen, Tao Wang, Weihao Yu et al.ICCV 2021 · 2,462 citations
- Attribute Prototype Network for Zero-Shot LearningWenjia Xu, Yongqin Xian, Jiuniu Wang, Bernt Schiele et al.NeurIPS 2020 · 392 citations
Related papers
- Progressive Semantic-Guided Vision Transformer for Zero-Shot LearningShiming Chen, Wenjin Hou, Salman H. Khan, Fahad Shahbaz KhanCVPR 2024
- TransZero: Attribute-Guided Transformer for Zero-Shot LearningShiming Chen, Ziming Hong, Yang Liu, Guo-Sen Xie et al.AAAI 2022 · 185 citations
- Rethinking Zero-Shot Learning: A Conditional Visual Classification PerspectiveKai Li, Martin Renqiang Min, Yun FuICCV 2019 · 151 citations
- Rethinking Generative Zero-Shot Learning: An Ensemble Learning Perspective for Recognising Visual PatchesZhi Chen, Sen Wang, Jingjing Li, Zi HuangACM MM 2020 · 33 citations
- Causal Visual-semantic Correlation for Zero-shot LearningShuhuang Chen, Dingjie Fu, Shiming Chen, Shuo Ye et al.ACM MM 2024 · 11 citations
