Visual Relationship Detection with Low Rank Non-Negative Tensor Decomposition
Mohammed Haroon Dupty, Zhen Zhang, Wee Sun Lee
摘要
We address the problem of Visual Relationship Detection (VRD) which aims to describe the relationships between pairs of objects in the form of triplets of (subject, predicate, object). We observe that given a pair of bounding box proposals, objects often participate in multiple relations implying the distribution of triplets is multimodal. We leverage the strong correlations within triplets to learn the joint distribution of triplet variables conditioned on the image and the bounding box proposals, doing away with the hitherto used independent distribution of triplets. To make learning the triplet joint distribution feasible, we introduce a novel technique of learning conditional triplet distributions in the form of their normalized low rank non-negative tensor decompositions. Normalized tensor decompositions take form of mixture distributions of discrete variables and thus are able to capture multimodality. This allows us to efficiently learn higher order discrete multimodal distributions and at the same time keep the parameter size manageable. We further model the probability of selecting an object proposal pair and include a relation triplet prior in our model. We show that each part of the model improves performance and the combination outperforms state-of-the-art score on the Visual Genome (VG) and Visual Relationship Detection (VRD) datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- HiLo: Exploiting High Low Frequency Relations for Unbiased Panoptic Scene Graph GenerationZijian Zhou, Miaojing Shi, Holger CaesarICCV 2023 · 被引用 29 次
- Grounding Consistency: Distilling Spatial Common Sense for Precise Visual Relationship DetectionMarkos Diomataris, Nikolaos Gkanatsios, Vassilis Pitsikalis, Petros MaragosICCV 2021 · 被引用 11 次
相关 Paper
- Localize, Assemble, and Predicate: Contextual Object Proposal Embedding for Visual Relation DetectionRuihai Wu, Kehan Xu, Chenchen Liu, Nan Zhuang 等AAAI 2020 · 被引用 7 次
- Video Visual Relation Detection via Iterative InferenceXindi Shang, Yicong Li, Junbin Xiao, Wei Ji 等ACM MM 2021 · 被引用 41 次
- Hierarchical Graph Attention Network for Visual Relationship DetectionLi Mi, Zhenzhong ChenCVPR 2020
- Structured Sparse R-CNN for Direct Scene Graph GenerationYao Teng, Limin WangCVPR 2022 · 被引用 66 次
- Visual Relationship Detection Using Part-and-Sum Transformers with Composite QueriesQi Dong, Zhuowen Tu, Haofu Liao, Yuting Zhang 等ICCV 2021 · 被引用 43 次
