Relational Proxies: Emergent Relationships as Fine-Grained Discriminators
Abhra Chaudhuri, Massimiliano Mancini, Zeynep Akata, Anjan Dutta
Abstract
Fine-grained categories that largely share the same set of parts cannot be discriminated based on part information alone, as they mostly differ in the way the local parts relate to the overall global structure of the object. We propose Relational Proxies, a novel approach that leverages the relational information between the global and local views of an object for encoding its semantic label. Starting with a rigorous formalization of the notion of distinguishability between fine-grained categories, we prove the necessary and sufficient conditions that a model must satisfy in order to learn the underlying decision boundaries in the fine-grained setting. We design Relational Proxies based on our theoretical findings and evaluate it on seven challenging fine-grained benchmark datasets and achieve state-of-the-art results on all of them, surpassing the performance of all existing works with a margin exceeding 4% in some cases. We also experimentally validate our theory on fine-grained distinguishability and obtain consistent results across multiple benchmarks. Implementation is available at https://github.com/abhrac/relational-proxies .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cf3570b8-4fd1-4181-ac4c-182aba402013Cited by top-tier papers1
Ask how each one uses itBuilds on13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Intriguing Properties of Vision TransformersMuzammal Naseer, Kanchana Ranasinghe, Salman Khan, Munawar Hayat et al.NeurIPS 2021 · 863 citations
- TransFG: A Transformer Architecture for Fine-Grained RecognitionJu He, Jieneng Chen, Shuai Liu, Adam Kortylewski et al.AAAI 2022 · 529 citations
- Learning Robust Representations via Multi-View Information BottleneckMarco Federici, Anjan Dutta, Patrick Forré, Nate Kushman et al.ICLR 2020 · 330 citations
Related papers
- Deep Discriminative Structure Proxy Hashing for Cross-modal RetrievalKun Cheng, Qibing Qin, Lei HuangICML 2026
- Unsupervised Part Discovery from Contrastive ReconstructionSubhabrata Choudhury, Iro Laina, Christian Rupprecht, Andrea VedaldiNeurIPS 2021 · 74 citations
- Part-level Semantic-guided Contrastive Learning for Fine-grained Visual ClassificationZhijian Lin, Hong HanICLR 2026
- Beyond the Attention: Distinguish the Discriminative and Confusable Features For Fine-grained Image ClassificationXiruo Shi, Liutong Xu, Pengfei Wang, Yuanyuan Gao et al.ACM MM 2020 · 12 citations
- Graph-Based High-Order Relation Discovery for Fine-Grained RecognitionYifan Zhao, Ke Yan, Feiyue Huang, Jia LiCVPR 2021
