Video Visual Relation Detection via Iterative Inference
Xindi Shang, Yicong Li, Junbin Xiao, Wei Ji, Tat-Seng Chua
摘要
The core problem of video visual relation detection (VidVRD) lies in accurately classifying the relation triplets, which comprise of the classes of subject and object entities, and the predicate classes of various relationships between them. Existing VidVRD approaches classify these three relation components in either independent or cascaded manner, thus fail to fully exploit the inter-dependency among them. In order to utilize this inter-dependency in tackling the challenges of visual relation recognition in videos, we propose a novel iterative relation inference approach for VidVRD. We derive our model from the viewpoint of joint relation classification which is light-weight yet effective, and propose a training approach to better learn the dependency knowledge from the likely correct triplet combinations. As such, the proposed inference approach is able to gradually refine each component based on its learnt dependency and the other two's predictions. Our ablation studies show that this iterative relation inference can empirically converge in a few steps and consistently boost the performance over baselines. Further, we incorporate it into a newly designed VidVRD architecture, named VidVRD-II (Iterative Inference), which generalizes well across different datasets. Experiments show that VidVRD-II achieves the start-of-the-art performance on both of ImageNet-VidVRD and VidOR benchmark datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Video Question Answering: Datasets, Algorithms and ChallengesYaoyao Zhong, Wei Ji, Junbin Xiao, Yicong Li 等EMNLP 2022 · 被引用 70 次
- Classification-Then-Grounding: Reformulating Video Scene Graphs as Temporal Bipartite GraphsKaifeng Gao, Long Chen, Yulei Niu, Jian Shao 等CVPR 2022 · 被引用 34 次
- Redundancy-aware Transformer for Video Question AnsweringYicong Li, Xun Yang, An Zhang, Chun Feng 等ACM MM 2023 · 被引用 23 次
- Partial Annotation-based Video Moment Retrieval via Iterative LearningWei Ji, Renjie Liang, Lizi Liao, Hao Fei 等ACM MM 2023 · 被引用 17 次
- VRDFormer: End-to-End Video Visual Relation Detection with TransformersSipeng Zheng, Shizhe Chen, Qin JinCVPR 2022 · 被引用 16 次
它引用的顶会 Paper12
- Deconfounded Video Moment Retrieval with Causal InterventionXun Yang, Fuli Feng, Wei Ji, Meng Wang 等SIGIR 2021 · 被引用 198 次
- Counterfactual Critic Multi-Agent Training for Scene Graph GenerationLong Chen, Hanwang Zhang, Jun Xiao, Xiangnan He 等ICCV 2019 · 被引用 165 次
- Detecting Unseen Visual Relations Using AnalogiesJulia Peyre, Josef Sivic, Ivan Laptev, Cordelia SchmidICCV 2019 · 被引用 135 次
- Tree-Augmented Cross-Modal Encoding for Complex-Query Video RetrievalXun Yang, Jianfeng Dong, Yixin Cao, Xun Wang 等SIGIR 2020 · 被引用 131 次
- Weakly-Supervised Video Object Grounding by Exploring Spatio-Temporal ContextsXun Yang, Xueliang Liu, Meng Jian, Xinjian Gao 等ACM MM 2020 · 被引用 47 次
相关 Paper
- Video Relation Detection via Multiple Hypothesis AssociationZixuan Su, Xindi Shang, Jingjing Chen, Yu-Gang Jiang 等ACM MM 2020 · 被引用 37 次
- Interventional Video Relation DetectionYicong Li, Xun Yang, Xindi Shang, Tat-Seng ChuaACM MM 2021 · 被引用 61 次
- Beyond Short-Term Snippet: Video Relation Detection With Spatio-Temporal Global ContextChenchen Liu, Yang Jin, Kehan Xu, Guoqiang Gong 等CVPR 2020
- VrdONE: One-stage Video Visual Relation DetectionXinjie Jiang, Chenxi Zheng, Xuemiao Xu, Bangzhen Liu 等ACM MM 2024 · 被引用 1 次
- Visual Relationship Detection with Low Rank Non-Negative Tensor DecompositionMohammed Haroon Dupty, Zhen Zhang, Wee Sun LeeAAAI 2020 · 被引用 9 次
