TVT: Three-Way Vision Transformer through Multi-Modal Hypersphere Learning for Zero-Shot Sketch-Based Image Retrieval
Jialin Tian, Xing Xu, Fumin Shen, Yang Yang, Heng Tao Shen
Abstract
In this paper, we study the zero-shot sketch-based image retrieval (ZS-SBIR) task, which retrieves natural images related to sketch queries from unseen categories. In the literature, convolutional neural networks (CNNs) have become the de-facto standard and they are either trained end-to-end or used to extract pre-trained features for images and sketches. However, CNNs are limited in modeling the global structural information of objects due to the intrinsic locality of convolution operations. To this end, we propose a Transformer-based approach called Three-Way Vision Transformer (TVT) to leverage the ability of Vision Transformer (ViT) to model global contexts due to the global self-attention mechanism. Going beyond simply applying ViT to this task, we propose a token-based strategy of adding fusion and distillation tokens and making them complementary to each other. Specifically, we integrate three ViTs, which are pre-trained on data of each modality, into a three-way pipeline through the processes of distillation and multi-modal hypersphere learning. The distillation process is proposed to supervise fusion ViT (ViT with an extra fusion token) with soft targets from modality-specific ViTs, which prevents fusion ViT from catastrophic forgetting. Furthermore, our method learns a multi-modal hypersphere by performing inter- and intra-modal alignment without loss of uniformity, which aims to bridge the modal gap between modalities of sketch and image and avoid the collapse in dimensions. Extensive experiments on three benchmark datasets, i.e., Sketchy, TU-Berlin, and QuickDraw, demonstrate the superiority of our TVT method over the state-of-the-art ZS-SBIR methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d170f252-e9f9-4c7f-91d6-620503d72a1aCited by top-tier papers5
- What can Discriminator do? Towards Box-free Ownership Verification of Generative Adversarial NetworksZiheng Huang, Boheng Li, Yan Cai, Run Wang et al.ICCV 2023 · 19 citations
- Zero-Shot Everything Sketch-Based Image Retrieval, and in Explainable StyleFengyin Lin, Mingkang Li, Da Li, Timothy M. Hospedales et al.CVPR 2023
- Modeling the Visual Ambiguity of Human SketchesYang Zhou, Ping Ni, Jin Wang, Senyun Jia et al.CVPR 2026
- Text-to-Image Diffusion Models are Great Sketch-Photo MatchmakersSubhadeep Koley, Ayan Kumar Bhunia, Aneeshan Sain, Pinaki Nath Chowdhury et al.CVPR 2024
- CLIP for All Things Zero-Shot Sketch-Based Image Retrieval, Fine-Grained or NotAneeshan Sain, Ayan Kumar Bhunia, Pinaki Nath Chowdhury, Subhadeep Koley et al.CVPR 2023
Builds on7
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 2,360 citations
- Semantic-Aware Knowledge Preservation for Zero-Shot Sketch-Based Image RetrievalQing Liu, Lingxi Xie, Huiyu Wang, Alan L. YuilleICCV 2019 · 126 citations
Related papers
- Prototype-based Selective Knowledge Distillation for Zero-Shot Sketch Based Image RetrievalKai Wang, Yifan Wang, Xing Xu, Xin Liu et al.ACM MM 2022 · 41 citations
- Relationship-Preserving Knowledge Distillation for Zero-Shot Sketch Based Image RetrievalJialin Tian, Xing Xu, Zheng Wang, Fumin Shen et al.ACM MM 2021 · 56 citations
- Structure-Aware Semantic-Aligned Network for Universal Cross-Domain RetrievalJialin Tian, Xing Xu, Kai Wang, Zuo Cao et al.SIGIR 2022 · 7 citations
- Asymmetric Mutual Alignment for Unsupervised Zero-Shot Sketch-Based Image RetrievalZhihui Yin, Jiexi Yan, Chenghao Xu, Cheng DengAAAI 2024 · 6 citations
- Progressive Semantic-Guided Vision Transformer for Zero-Shot LearningShiming Chen, Wenjin Hou, Salman H. Khan, Fahad Shahbaz KhanCVPR 2024
