Single Domain Generalization for Few-Shot Counting via Universal Representation Matching
Xianing Chen, Si Huo, Borui Jiang, Hailin Hu, Xinghao Chen
Abstract
Few-shot counting estimates the number of target objects in an image using only a few annotated exemplars. However, domain shift severely hinders existing methods to generalize to unseen scenarios. This falls into the realm of single domain generalization that remains unexplored in fewshot counting. To solve this problem, we begin by analyzing the main limitations of current methods, which typically follow a standard pipeline that extract the object prototypes from exemplars and then match them with image feature to construct the correlation map. We argue that existing methods overlook the significance of learning highly generalized prototypes. Building on this insight, we propose the first single domain generalization few-shot counting model, Universal Representation Matching, termed URM. Our primary contribution is the discovery that incorporating universal vision-language representations distilled from a large scale pretrained vision-language model into the correlation construction process substantially improves robustness to domain shifts without compromising in domain performance. As a result, URM achieves state-of-the-art performance on both in domain and the newly introduced domain generalization setting. * Corresponding authors. ** Code is available at https://github.com/jbr97/URM . (a) domain generalization for few-shot counting (b) extract-then-match (c) universal representation matching
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a8390b60-6c8f-4a1d-a74b-db39f449872cCited by top-tier papers1
Ask how each one uses itBuilds on38
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Decoupled Knowledge DistillationBorui Zhao, Quan Cui, Renjie Song, Yiyu Qiu et al.CVPR 2022 · 835 citations
- What matters when building vision-language models?Hugo Laurençon, Léo Tronchon, Matthieu Cord, Victor SanhNeurIPS 2024 · 401 citations
Related papers
- Universal Representation Learning from Multiple Domains for Few-shot ClassificationWei-Hong Li, Xialei Liu, Hakan BilenICCV 2021 · 114 citations
- Universal Few-shot Learning of Dense Prediction Tasks with Visual Token MatchingDonggyun Kim, Jinwoo Kim, Seongwoong Cho, Chong Luo et al.ICLR 2023 · 3 citations
- A Universal Representation Transformer Layer for Few-Shot Image ClassificationLu Liu, William L. Hamilton, Guodong Long, Jing Jiang et al.ICLR 2021 · 143 citations
- A Novel Unified Architecture for Low-Shot Counting by Detection and SegmentationJer Pelhan, Alan Lukezic, Vitjan Zavrtanik, Matej KristanNeurIPS 2024 · 29 citations
- Learning To Count EverythingViresh Ranjan, Udbhav Sharma, Thu Nguyen, Minh HoaiCVPR 2021
