Interventional Video Relation Detection
Yicong Li, Xun Yang, Xindi Shang, Tat-Seng Chua
Abstract
Video Visual Relation Detection (VidVRD) aims to semantically describe the dynamic interactions across visual concepts localized in a video in the form of subject, predicate, object. It can help to mitigate the semantic gap between vision and language in video understanding, thus receiving increasing attention in multimedia communities. Existing efforts primarily leverage the multimodal/spatio-temporal feature fusion to augment the representation of object trajectories as well as their interactions and formulate the prediction of predicates as a multi-class classification task. Despite their effectiveness, existing models ignore the severe long-tailed bias in VidVRD datasets. As a result, the models' prediction will be easily biased towards the popular head predicates (e.g., next-to and in-front-of), thus leading to poor generalizability.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 5ab3d664-031f-4621-9628-3e77bfb7c514Cited by top-tier papers23
- Invariant Grounding for Video Question AnsweringYicong Li, Xiang Wang, Junbin Xiao, Wei Ji et al.CVPR 2022 · 108 citations
- Counterfactual Reasoning for Out-of-distribution Multimodal Sentiment AnalysisTeng Sun, Wenjie Wang, Liqiang Jing, Yiran Cui et al.ACM MM 2022 · 65 citations
- EI-CLIP: Entity-aware Interventional Contrastive Learning for E-commerce Cross-modal RetrievalHaoyu Ma, Handong Zhao, Zhe Lin, Ajinkya Kale et al.CVPR 2022 · 56 citations
- Causality-Inspired Invariant Representation Learning for Text-Based Person RetrievalYu Liu, Guihe Qin, Haipeng Chen, Zhiyong Cheng et al.AAAI 2024 · 44 citations
- Uncovering Main Causalities for Long-tailed Information ExtractionGuoshun Nan, Jiaqi Zeng, Rui Qiao, Zhijiang Guo et al.EMNLP 2021 · 39 citations
Related papers
- Video Relation Detection via Multiple Hypothesis AssociationZixuan Su, Xindi Shang, Jingjing Chen, Yu-Gang Jiang et al.ACM MM 2020 · 37 citations
- Beyond Short-Term Snippet: Video Relation Detection With Spatio-Temporal Global ContextChenchen Liu, Yang Jin, Kehan Xu, Guoqiang Gong et al.CVPR 2020
- Video Visual Relation Detection via Iterative InferenceXindi Shang, Yicong Li, Junbin Xiao, Wei Ji et al.ACM MM 2021 · 41 citations
- VrdONE: One-stage Video Visual Relation DetectionXinjie Jiang, Chenxi Zheng, Xuemiao Xu, Bangzhen Liu et al.ACM MM 2024 · 1 citation
- VRDFormer: End-to-End Video Visual Relation Detection with TransformersSipeng Zheng, Shizhe Chen, Qin JinCVPR 2022 · 16 citations
