Interventional Video Relation Detection
Yicong Li, Xun Yang, Xindi Shang, Tat-Seng Chua
摘要
Video Visual Relation Detection (VidVRD) aims to semantically describe the dynamic interactions across visual concepts localized in a video in the form of subject, predicate, object. It can help to mitigate the semantic gap between vision and language in video understanding, thus receiving increasing attention in multimedia communities. Existing efforts primarily leverage the multimodal/spatio-temporal feature fusion to augment the representation of object trajectories as well as their interactions and formulate the prediction of predicates as a multi-class classification task. Despite their effectiveness, existing models ignore the severe long-tailed bias in VidVRD datasets. As a result, the models' prediction will be easily biased towards the popular head predicates (e.g., next-to and in-front-of), thus leading to poor generalizability.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper23
- Invariant Grounding for Video Question AnsweringYicong Li, Xiang Wang, Junbin Xiao, Wei Ji 等CVPR 2022 · 被引用 108 次
- Counterfactual Reasoning for Out-of-distribution Multimodal Sentiment AnalysisTeng Sun, Wenjie Wang, Liqiang Jing, Yiran Cui 等ACM MM 2022 · 被引用 65 次
- EI-CLIP: Entity-aware Interventional Contrastive Learning for E-commerce Cross-modal RetrievalHaoyu Ma, Handong Zhao, Zhe Lin, Ajinkya Kale 等CVPR 2022 · 被引用 56 次
- Causality-Inspired Invariant Representation Learning for Text-Based Person RetrievalYu Liu, Guihe Qin, Haipeng Chen, Zhiyong Cheng 等AAAI 2024 · 被引用 44 次
- Uncovering Main Causalities for Long-tailed Information ExtractionGuoshun Nan, Jiaqi Zeng, Rui Qiao, Zhijiang Guo 等EMNLP 2021 · 被引用 39 次
相关 Paper
- Video Relation Detection via Multiple Hypothesis AssociationZixuan Su, Xindi Shang, Jingjing Chen, Yu-Gang Jiang 等ACM MM 2020 · 被引用 37 次
- Beyond Short-Term Snippet: Video Relation Detection With Spatio-Temporal Global ContextChenchen Liu, Yang Jin, Kehan Xu, Guoqiang Gong 等CVPR 2020
- Video Visual Relation Detection via Iterative InferenceXindi Shang, Yicong Li, Junbin Xiao, Wei Ji 等ACM MM 2021 · 被引用 41 次
- VrdONE: One-stage Video Visual Relation DetectionXinjie Jiang, Chenxi Zheng, Xuemiao Xu, Bangzhen Liu 等ACM MM 2024 · 被引用 1 次
- VRDFormer: End-to-End Video Visual Relation Detection with TransformersSipeng Zheng, Shizhe Chen, Qin JinCVPR 2022 · 被引用 16 次
