Visual Relationship Detection with Low Rank Non-Negative Tensor Decomposition
Mohammed Haroon Dupty, Zhen Zhang, Wee Sun Lee
Abstract
We address the problem of Visual Relationship Detection (VRD) which aims to describe the relationships between pairs of objects in the form of triplets of (subject, predicate, object). We observe that given a pair of bounding box proposals, objects often participate in multiple relations implying the distribution of triplets is multimodal. We leverage the strong correlations within triplets to learn the joint distribution of triplet variables conditioned on the image and the bounding box proposals, doing away with the hitherto used independent distribution of triplets. To make learning the triplet joint distribution feasible, we introduce a novel technique of learning conditional triplet distributions in the form of their normalized low rank non-negative tensor decompositions. Normalized tensor decompositions take form of mixture distributions of discrete variables and thus are able to capture multimodality. This allows us to efficiently learn higher order discrete multimodal distributions and at the same time keep the parameter size manageable. We further model the probability of selecting an object proposal pair and include a relation triplet prior in our model. We show that each part of the model improves performance and the combination outperforms state-of-the-art score on the Visual Genome (VG) and Visual Relationship Detection (VRD) datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3a5dbddf-a787-42b2-a989-8fef1592faf1Cited by top-tier papers2
- HiLo: Exploiting High Low Frequency Relations for Unbiased Panoptic Scene Graph GenerationZijian Zhou, Miaojing Shi, Holger CaesarICCV 2023 · 29 citations
- Grounding Consistency: Distilling Spatial Common Sense for Precise Visual Relationship DetectionMarkos Diomataris, Nikolaos Gkanatsios, Vassilis Pitsikalis, Petros MaragosICCV 2021 · 11 citations
Related papers
- Localize, Assemble, and Predicate: Contextual Object Proposal Embedding for Visual Relation DetectionRuihai Wu, Kehan Xu, Chenchen Liu, Nan Zhuang et al.AAAI 2020 · 7 citations
- Video Visual Relation Detection via Iterative InferenceXindi Shang, Yicong Li, Junbin Xiao, Wei Ji et al.ACM MM 2021 · 41 citations
- Hierarchical Graph Attention Network for Visual Relationship DetectionLi Mi, Zhenzhong ChenCVPR 2020
- Structured Sparse R-CNN for Direct Scene Graph GenerationYao Teng, Limin WangCVPR 2022 · 66 citations
- Visual Relationship Detection Using Part-and-Sum Transformers with Composite QueriesQi Dong, Zhuowen Tu, Haofu Liao, Yuting Zhang et al.ICCV 2021 · 43 citations
