CrossTransformers: spatially-aware few-shot transfer
Carl Doersch, Ankush Gupta, Andrew Zisserman
Abstract
Given new tasks with very little data-such as new classes in a classification problem or a domain shift in the input-performance of modern vision systems degrades remarkably quickly. In this work, we illustrate how the neural network representations which underpin modern vision systems are subject to supervision collapse, whereby they lose any information that is not necessary for performing the training task, including information that may be necessary for transfer to new tasks or domains. We then propose two methods to mitigate this problem. First, we employ self-supervised learning to encourage general-purpose features that transfer better. Second, we propose a novel Transformer based neural network architecture called CrossTransformers, which can take a small number of labeled images and an unlabeled query, find coarse spatial correspondence between the query and the labeled images, and then infer class membership by computing distances between spatially-corresponding features. The result is a classifier that is more robust to task and domain shift, which we demonstrate via state-of-theart performance on Meta-Dataset, a recent dataset for evaluating transfer from ImageNet to many other vision datasets. Code and pretrained checkpoints available at: https://github.com/google-research/meta-dataset .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5dcaa42c-702a-424f-90c5-fac993d4abaaCited by top-tier papers85
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Segmenter: Transformer for Semantic SegmentationRobin Strudel, Ricardo Garcia, Ivan Laptev, Cordelia SchmidICCV 2021 · 1,898 citations
- Keeping Your Eye on the Ball: Trajectory Attention in Video TransformersMandela Patrick, Dylan Campbell, Yuki M. Asano, Ishan Misra et al.NeurIPS 2021 · 382 citations
- Joint Distribution Matters: Deep Brownian Distance Covariance for Few-Shot ClassificationJiangtao Xie, Fei Long, Jiaming Lv, Qilong Wang et al.CVPR 2022 · 270 citations
- Relational Embedding for Few-Shot ClassificationDahyun Kang, Heeseung Kwon, Juhong Min, Minsu ChoICCV 2021 · 254 citations
Builds on7
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Meta-Dataset: A Dataset of Datasets for Learning to Learn from Few ExamplesEleni Triantafillou, Tyler Zhu, Vincent Dumoulin, Pascal Lamblin et al.ICLR 2020 · 692 citations
- A Baseline for Few-Shot Image ClassificationGuneet Singh Dhillon, Pratik Chaudhari, Avinash Ravichandran, Stefano SoattoICLR 2020 · 640 citations
- Boosting Few-Shot Visual Learning With Self-SupervisionSpyros Gidaris, Andrei Bursuc, Nikos Komodakis, Patrick Pérez et al.ICCV 2019 · 445 citations
- Learning Compositional Representations for Few-Shot RecognitionPavel Tokmakov, Yu-Xiong Wang, Martial HebertICCV 2019 · 133 citations
Related papers
- Efficient Training of Visual Transformers with Small DatasetsYahui Liu, Enver Sangineto, Wei Bi, Nicu Sebe et al.NeurIPS 2021 · 238 citations
- A Universal Representation Transformer Layer for Few-Shot Image ClassificationLu Liu, William L. Hamilton, Guodong Long, Jing Jiang et al.ICLR 2021 · 143 citations
- Pushing the Limits of Simple Pipelines for Few-Shot Learning: External Data and Fine-Tuning Make a DifferenceShell Xu Hu, Da Li, Jan Stühmer, Minyoung Kim et al.CVPR 2022 · 161 citations
- Efficient Self-supervised Vision Transformers for Representation LearningChunyuan Li, Jianwei Yang, Pengchuan Zhang, Mei Gao et al.ICLR 2022 · 228 citations
- Probabilistic Self-supervised Representation Learning via Scoring Rules MinimizationAmirhossein Vahidi, Simon Schoßer, Lisa Wimmer, Yawei Li et al.ICLR 2024 · 6 citations
