One-Shot Learning for Long-Tail Visual Relation Detection
Weitao Wang, Meng Wang, Sen Wang, Guodong Long, Lina Yao, Guilin Qi, Yang Chen
Abstract
The aim of visual relation detection is to provide a comprehensive understanding of an image by describing all the objects within the scene, and how they relate to each other, in < object-predicate-object > form; for example, < person-lean on-wall > . This ability is vital for image captioning, visual question answering, and many other applications. However, visual relationships have long-tailed distributions and, thus, the limited availability of training samples is hampering the practicability of conventional detection approaches. With this in mind, we designed a novel model for visual relation detection that works in one-shot settings. The embeddings of objects and predicates are extracted through a network that includes a feature-level attention mechanism. Attention alleviates some of the problems with feature sparsity, and the resulting representations capture more discriminative latent features. The core of our model is a dual graph neural network that passes and aggregates the context information of predicates and objects in an episodic training scheme to improve recognition of the one-shot predicates and then generate the triplets. To the best of our knowledge, we are the first to center on the viability of one-shot learning for visual relation detection. Extensive experiments on two newly-constructed datasets show that our model significantly improved the performance of two tasks PredCls and SGCls from 2.8% to 12.2% compared with state-of-the-art baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 026a6405-3132-4668-9801-4c97d43fc2faCited by top-tier papers2
- Grounding Consistency: Distilling Spatial Common Sense for Precise Visual Relationship DetectionMarkos Diomataris, Nikolaos Gkanatsios, Vassilis Pitsikalis, Petros MaragosICCV 2021 · 11 citations
- Few-Shot Referring Relationships in VideosYogesh Kumar, Anand MishraCVPR 2023
Related papers
- One-shot Scene Graph GenerationYuyu Guo, Jingkuan Song, Lianli Gao, Heng Tao ShenACM MM 2020 · 26 citations
- Hierarchical Graph Attention Network for Visual Relationship DetectionLi Mi, Zhenzhong ChenCVPR 2020
- Leveraging Predicate and Triplet Learning for Scene Graph GenerationJiankai Li, Yunhong Wang, Xiefan Guo, Ruijie Yang et al.CVPR 2024
- Memory-Based Network for Scene Graph with Unbalanced RelationsWeitao Wang, Ruyang Liu, Meng Wang, Sen Wang et al.ACM MM 2020 · 11 citations
- Localize, Assemble, and Predicate: Contextual Object Proposal Embedding for Visual Relation DetectionRuihai Wu, Kehan Xu, Chenchen Liu, Nan Zhuang et al.AAAI 2020 · 7 citations
