CrossTransformers: spatially-aware few-shot transfer
Carl Doersch, Ankush Gupta, Andrew Zisserman
摘要
Given new tasks with very little data-such as new classes in a classification problem or a domain shift in the input-performance of modern vision systems degrades remarkably quickly. In this work, we illustrate how the neural network representations which underpin modern vision systems are subject to supervision collapse, whereby they lose any information that is not necessary for performing the training task, including information that may be necessary for transfer to new tasks or domains. We then propose two methods to mitigate this problem. First, we employ self-supervised learning to encourage general-purpose features that transfer better. Second, we propose a novel Transformer based neural network architecture called CrossTransformers, which can take a small number of labeled images and an unlabeled query, find coarse spatial correspondence between the query and the labeled images, and then infer class membership by computing distances between spatially-corresponding features. The result is a classifier that is more robust to task and domain shift, which we demonstrate via state-of-theart performance on Meta-Dataset, a recent dataset for evaluating transfer from ImageNet to many other vision datasets. Code and pretrained checkpoints available at: https://github.com/google-research/meta-dataset .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper85
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- Segmenter: Transformer for Semantic SegmentationRobin Strudel, Ricardo Garcia, Ivan Laptev, Cordelia SchmidICCV 2021 · 被引用 1,898 次
- Keeping Your Eye on the Ball: Trajectory Attention in Video TransformersMandela Patrick, Dylan Campbell, Yuki M. Asano, Ishan Misra 等NeurIPS 2021 · 被引用 382 次
- Joint Distribution Matters: Deep Brownian Distance Covariance for Few-Shot ClassificationJiangtao Xie, Fei Long, Jiaming Lv, Qilong Wang 等CVPR 2022 · 被引用 270 次
- Relational Embedding for Few-Shot ClassificationDahyun Kang, Heeseung Kwon, Juhong Min, Minsu ChoICCV 2021 · 被引用 254 次
它引用的顶会 Paper7
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Meta-Dataset: A Dataset of Datasets for Learning to Learn from Few ExamplesEleni Triantafillou, Tyler Zhu, Vincent Dumoulin, Pascal Lamblin 等ICLR 2020 · 被引用 692 次
- A Baseline for Few-Shot Image ClassificationGuneet Singh Dhillon, Pratik Chaudhari, Avinash Ravichandran, Stefano SoattoICLR 2020 · 被引用 640 次
- Boosting Few-Shot Visual Learning With Self-SupervisionSpyros Gidaris, Andrei Bursuc, Nikos Komodakis, Patrick Pérez 等ICCV 2019 · 被引用 445 次
- Learning Compositional Representations for Few-Shot RecognitionPavel Tokmakov, Yu-Xiong Wang, Martial HebertICCV 2019 · 被引用 133 次
相关 Paper
- Efficient Training of Visual Transformers with Small DatasetsYahui Liu, Enver Sangineto, Wei Bi, Nicu Sebe 等NeurIPS 2021 · 被引用 238 次
- A Universal Representation Transformer Layer for Few-Shot Image ClassificationLu Liu, William L. Hamilton, Guodong Long, Jing Jiang 等ICLR 2021 · 被引用 143 次
- Pushing the Limits of Simple Pipelines for Few-Shot Learning: External Data and Fine-Tuning Make a DifferenceShell Xu Hu, Da Li, Jan Stühmer, Minyoung Kim 等CVPR 2022 · 被引用 161 次
- Efficient Self-supervised Vision Transformers for Representation LearningChunyuan Li, Jianwei Yang, Pengchuan Zhang, Mei Gao 等ICLR 2022 · 被引用 228 次
- Probabilistic Self-supervised Representation Learning via Scoring Rules MinimizationAmirhossein Vahidi, Simon Schoßer, Lisa Wimmer, Yawei Li 等ICLR 2024 · 被引用 6 次
