Retrieval Augmented Classification for Long-Tail Visual Recognition
Alexander Long, Wei Yin, Thalaiyasingam Ajanthan, Vu Nguyen, Pulak Purkait, Ravi Garg, Alan Blair, Chunhua Shen, Anton van den Hengel
摘要
We introduce Retrieval Augmented Classification (RAC), a generic approach to augmenting standard image classification pipelines with an explicit retrieval module. RAC consists of a standard base image encoder fused with a parallel retrieval branch that queries a non-parametric external memory of pre-encoded images and associated text snippets. We apply RAC to the problem of long-tail classification and demonstrate a significant improvement over previous state-of-the-art on Places365-LT and iNaturalist-2018 (14.5% and 6.7% respectively), despite using only the training datasets themselves as the external information source. We demonstrate that RAC's retrieval module, without prompting, learns a high level of accuracy on tail classes. This, in turn, frees the base encoder to focus on common classes, and improve its performance thereon. RAC represents an alternative approach to utilizing large, pretrained models without requiring fine-tuning, as well as a first step towards more effectively making use of external memory within common computer vision architectures.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper46
- Long-Tail Learning with Foundation Model: Heavy Fine-Tuning HurtsJiang-Xin Shi, Tong Wei, Zhi Zhou, Jie-Jing Shao 等ICML 2024 · 被引用 78 次
- TabR: Tabular Deep Learning Meets Nearest NeighborsYury Gorishniy, Ivan Rubachev, Nikolay Kartashev, Daniil Shlenskii 等ICLR 2024 · 被引用 78 次
- Retrieval-Enhanced Contrastive Vision-Text ModelsAhmet Iscen, Mathilde Caron, Alireza Fathi, Cordelia SchmidICLR 2024 · 被引用 44 次
- Re-Imagen: Retrieval-Augmented Text-to-Image GeneratorWenhu Chen, Hexiang Hu, Chitwan Saharia, William W. CohenICLR 2023 · 被引用 44 次
- Label-Retrieval-Augmented Diffusion Models for Learning from Noisy LabelsJian Chen, Ruiyi Zhang, Tong Yu, Rohan Sharma 等NeurIPS 2023 · 被引用 44 次
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Decoupling Representation and Classifier for Long-Tailed RecognitionBingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan 等ICLR 2020 · 被引用 1,496 次
- Generalization through Memorization: Nearest Neighbor Language ModelsUrvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer 等ICLR 2020 · 被引用 1,038 次
相关 Paper
- Improving Image Recognition by Retrieving from Web-Scale Image-Text DataAhmet Iscen, Alireza Fathi, Cordelia SchmidCVPR 2023
- MuRAG: Multimodal Retrieval-Augmented Generator for Open Question Answering over Images and TextWenhu Chen, Hexiang Hu, Xi Chen, Pat Verga 等EMNLP 2022 · 被引用 89 次
- Inflated Episodic Memory With Region Self-Attention for Long-Tailed Visual RecognitionLinchao Zhu, Yi YangCVPR 2020
- HGLTR: Hierarchical Knowledge Injection for Calibrating Pre-trained Models in Long-Tail RecognitionJinpeng Zheng, Shao-Yuan Li, Gan Xu, Wenhai Wan 等AAAI 2026
- Evcap: Retrieval-Augmented Image Captioning with External Visual-Name Memory for Open-World ComprehensionJiaxuan Li, Duc Minh Vo, Akihiro Sugimoto, Hideki NakayamaCVPR 2024
