Single Domain Generalization for Few-Shot Counting via Universal Representation Matching
Xianing Chen, Si Huo, Borui Jiang, Hailin Hu, Xinghao Chen
摘要
Few-shot counting estimates the number of target objects in an image using only a few annotated exemplars. However, domain shift severely hinders existing methods to generalize to unseen scenarios. This falls into the realm of single domain generalization that remains unexplored in fewshot counting. To solve this problem, we begin by analyzing the main limitations of current methods, which typically follow a standard pipeline that extract the object prototypes from exemplars and then match them with image feature to construct the correlation map. We argue that existing methods overlook the significance of learning highly generalized prototypes. Building on this insight, we propose the first single domain generalization few-shot counting model, Universal Representation Matching, termed URM. Our primary contribution is the discovery that incorporating universal vision-language representations distilled from a large scale pretrained vision-language model into the correlation construction process substantially improves robustness to domain shifts without compromising in domain performance. As a result, URM achieves state-of-the-art performance on both in domain and the newly introduced domain generalization setting. * Corresponding authors. ** Code is available at https://github.com/jbr97/URM . (a) domain generalization for few-shot counting (b) extract-then-match (c) universal representation matching
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper38
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- Decoupled Knowledge DistillationBorui Zhao, Quan Cui, Renjie Song, Yiyu Qiu 等CVPR 2022 · 被引用 835 次
- What matters when building vision-language models?Hugo Laurençon, Léo Tronchon, Matthieu Cord, Victor SanhNeurIPS 2024 · 被引用 401 次
相关 Paper
- Universal Representation Learning from Multiple Domains for Few-shot ClassificationWei-Hong Li, Xialei Liu, Hakan BilenICCV 2021 · 被引用 114 次
- Universal Few-shot Learning of Dense Prediction Tasks with Visual Token MatchingDonggyun Kim, Jinwoo Kim, Seongwoong Cho, Chong Luo 等ICLR 2023 · 被引用 3 次
- A Universal Representation Transformer Layer for Few-Shot Image ClassificationLu Liu, William L. Hamilton, Guodong Long, Jing Jiang 等ICLR 2021 · 被引用 143 次
- A Novel Unified Architecture for Low-Shot Counting by Detection and SegmentationJer Pelhan, Alan Lukezic, Vitjan Zavrtanik, Matej KristanNeurIPS 2024 · 被引用 29 次
- Learning To Count EverythingViresh Ranjan, Udbhav Sharma, Thu Nguyen, Minh HoaiCVPR 2021
