Few-Shot Visual Relationship Co-Localization
Revant Teotia, Vaibhav Mishra, Mayank Maheshwari, Anand Mishra
摘要
In this paper, given a small bag of images, each containing a common but latent predicate, we are interested in localizing visual subject-object pairs connected via the common predicate in each of the images. We refer to this novel problem as visual relationship co-localization or VRC as an abbreviation. VRC is a challenging task, even more so than the well-studied object co-localization task. This becomes further challenging when using just a few images, the model has to learn to co-localize visual subject-object pairs connected via unseen predicates. To solve VRC, we propose an optimization framework to select a common visual relationship in each image of the bag. The goal of the optimization framework is to find the optimal solution by learning visual relationship similarity across images in a few-shot setting. To obtain robust visual relationship representation, we utilize a simple yet effective technique that learns relationship embedding as a translation vector from visual subject to visual object in a shared space. Further, to learn visual relationship similarity, we utilize a proven meta-learning technique commonly used for few-shot classification tasks. Finally, to tackle the combinatorial complexity challenge arising from an exponential number of feasible solutions, we use a greedy approximation inference algorithm that selects approximately the best solution. We extensively evaluate our proposed framework on variations of bag sizes obtained from two challenging public datasets, namely VrR-VG and VG-150, and achieve impressive visual co-localization performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper4
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- VrR-VG: Refocusing Visually-Relevant RelationshipsYuanzhi Liang, Yalong Bai, Wei Zhang, Xueming Qian 等ICCV 2019 · 被引用 93 次
- SILCO: Show a Few Images, Localize the Common ObjectTao Hu, Pascal Mettes, Jia-Hong Huang, Cees SnoekICCV 2019 · 被引用 49 次
- Learning to Find Common Objects Across Few Image CollectionsAmirreza Shaban, Amir Rahimi, Shray Bansal, Stephen Gould 等ICCV 2019 · 被引用 9 次
相关 Paper
- Meta-RCNN: Meta Learning for Few-Shot Object DetectionXiongwei Wu, Doyen Sahoo, Steven C. H. HoiACM MM 2020 · 被引用 94 次
- Meta Learning to Bridge Vision and Language Models for Multimodal Few-Shot LearningIvona Najdenkoska, Xiantong Zhen, Marcel WorringICLR 2023 · 被引用 8 次
- Detecting Unseen Visual Relations Using AnalogiesJulia Peyre, Josef Sivic, Ivan Laptev, Cordelia SchmidICCV 2019 · 被引用 135 次
- Semantic Relation Reasoning for Shot-Stable Few-Shot Object DetectionChenchen Zhu, Fangyi Chen, Uzair Ahmed, Zhiqiang Shen 等CVPR 2021
- Task Cooperation for Semi-Supervised Few-Shot LearningHan-Jia Ye, Xin-Chun Li, De-Chuan ZhanAAAI 2021 · 被引用 20 次
