Few-Shot Visual Relationship Co-Localization
Revant Teotia, Vaibhav Mishra, Mayank Maheshwari, Anand Mishra
Abstract
In this paper, given a small bag of images, each containing a common but latent predicate, we are interested in localizing visual subject-object pairs connected via the common predicate in each of the images. We refer to this novel problem as visual relationship co-localization or VRC as an abbreviation. VRC is a challenging task, even more so than the well-studied object co-localization task. This becomes further challenging when using just a few images, the model has to learn to co-localize visual subject-object pairs connected via unseen predicates. To solve VRC, we propose an optimization framework to select a common visual relationship in each image of the bag. The goal of the optimization framework is to find the optimal solution by learning visual relationship similarity across images in a few-shot setting. To obtain robust visual relationship representation, we utilize a simple yet effective technique that learns relationship embedding as a translation vector from visual subject to visual object in a shared space. Further, to learn visual relationship similarity, we utilize a proven meta-learning technique commonly used for few-shot classification tasks. Finally, to tackle the combinatorial complexity challenge arising from an exponential number of feasible solutions, we use a greedy approximation inference algorithm that selects approximately the best solution. We extensively evaluate our proposed framework on variations of bag sizes obtained from two challenging public datasets, namely VrR-VG and VG-150, and achieve impressive visual co-localization performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9809f9bd-77e6-42d6-827b-1c55017ebcfaCited by top-tier papers1
Ask how each one uses itBuilds on4
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- VrR-VG: Refocusing Visually-Relevant RelationshipsYuanzhi Liang, Yalong Bai, Wei Zhang, Xueming Qian et al.ICCV 2019 · 93 citations
- SILCO: Show a Few Images, Localize the Common ObjectTao Hu, Pascal Mettes, Jia-Hong Huang, Cees SnoekICCV 2019 · 49 citations
- Learning to Find Common Objects Across Few Image CollectionsAmirreza Shaban, Amir Rahimi, Shray Bansal, Stephen Gould et al.ICCV 2019 · 9 citations
Related papers
- Meta-RCNN: Meta Learning for Few-Shot Object DetectionXiongwei Wu, Doyen Sahoo, Steven C. H. HoiACM MM 2020 · 94 citations
- Meta Learning to Bridge Vision and Language Models for Multimodal Few-Shot LearningIvona Najdenkoska, Xiantong Zhen, Marcel WorringICLR 2023 · 8 citations
- Detecting Unseen Visual Relations Using AnalogiesJulia Peyre, Josef Sivic, Ivan Laptev, Cordelia SchmidICCV 2019 · 135 citations
- Semantic Relation Reasoning for Shot-Stable Few-Shot Object DetectionChenchen Zhu, Fangyi Chen, Uzair Ahmed, Zhiqiang Shen et al.CVPR 2021
- Task Cooperation for Semi-Supervised Few-Shot LearningHan-Jia Ye, Xin-Chun Li, De-Chuan ZhanAAAI 2021 · 20 citations
