TVT: Three-Way Vision Transformer through Multi-Modal Hypersphere Learning for Zero-Shot Sketch-Based Image Retrieval
Jialin Tian, Xing Xu, Fumin Shen, Yang Yang, Heng Tao Shen
摘要
In this paper, we study the zero-shot sketch-based image retrieval (ZS-SBIR) task, which retrieves natural images related to sketch queries from unseen categories. In the literature, convolutional neural networks (CNNs) have become the de-facto standard and they are either trained end-to-end or used to extract pre-trained features for images and sketches. However, CNNs are limited in modeling the global structural information of objects due to the intrinsic locality of convolution operations. To this end, we propose a Transformer-based approach called Three-Way Vision Transformer (TVT) to leverage the ability of Vision Transformer (ViT) to model global contexts due to the global self-attention mechanism. Going beyond simply applying ViT to this task, we propose a token-based strategy of adding fusion and distillation tokens and making them complementary to each other. Specifically, we integrate three ViTs, which are pre-trained on data of each modality, into a three-way pipeline through the processes of distillation and multi-modal hypersphere learning. The distillation process is proposed to supervise fusion ViT (ViT with an extra fusion token) with soft targets from modality-specific ViTs, which prevents fusion ViT from catastrophic forgetting. Furthermore, our method learns a multi-modal hypersphere by performing inter- and intra-modal alignment without loss of uniformity, which aims to bridge the modal gap between modalities of sketch and image and avoid the collapse in dimensions. Extensive experiments on three benchmark datasets, i.e., Sketchy, TU-Berlin, and QuickDraw, demonstrate the superiority of our TVT method over the state-of-the-art ZS-SBIR methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- What can Discriminator do? Towards Box-free Ownership Verification of Generative Adversarial NetworksZiheng Huang, Boheng Li, Yan Cai, Run Wang 等ICCV 2023 · 被引用 19 次
- Zero-Shot Everything Sketch-Based Image Retrieval, and in Explainable StyleFengyin Lin, Mingkang Li, Da Li, Timothy M. Hospedales 等CVPR 2023
- Modeling the Visual Ambiguity of Human SketchesYang Zhou, Ping Ni, Jin Wang, Senyun Jia 等CVPR 2026
- Text-to-Image Diffusion Models are Great Sketch-Photo MatchmakersSubhadeep Koley, Ayan Kumar Bhunia, Aneeshan Sain, Pinaki Nath Chowdhury 等CVPR 2024
- CLIP for All Things Zero-Shot Sketch-Based Image Retrieval, Fine-Grained or NotAneeshan Sain, Ayan Kumar Bhunia, Pinaki Nath Chowdhury, Subhadeep Koley 等CVPR 2023
它引用的顶会 Paper7
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 被引用 2,360 次
- Semantic-Aware Knowledge Preservation for Zero-Shot Sketch-Based Image RetrievalQing Liu, Lingxi Xie, Huiyu Wang, Alan L. YuilleICCV 2019 · 被引用 126 次
相关 Paper
- Prototype-based Selective Knowledge Distillation for Zero-Shot Sketch Based Image RetrievalKai Wang, Yifan Wang, Xing Xu, Xin Liu 等ACM MM 2022 · 被引用 41 次
- Relationship-Preserving Knowledge Distillation for Zero-Shot Sketch Based Image RetrievalJialin Tian, Xing Xu, Zheng Wang, Fumin Shen 等ACM MM 2021 · 被引用 56 次
- Structure-Aware Semantic-Aligned Network for Universal Cross-Domain RetrievalJialin Tian, Xing Xu, Kai Wang, Zuo Cao 等SIGIR 2022 · 被引用 7 次
- Asymmetric Mutual Alignment for Unsupervised Zero-Shot Sketch-Based Image RetrievalZhihui Yin, Jiexi Yan, Chenghao Xu, Cheng DengAAAI 2024 · 被引用 6 次
- Progressive Semantic-Guided Vision Transformer for Zero-Shot LearningShiming Chen, Wenjin Hou, Salman H. Khan, Fahad Shahbaz KhanCVPR 2024
