Beyond Short-Term Snippet: Video Relation Detection With Spatio-Temporal Global Context
Chenchen Liu, Yang Jin, Kehan Xu, Guoqiang Gong, Yadong Mu
摘要
Video visual relation detection (VidVRD) aims to describe all interacting objects in a video. Different from relationships in static images, videos contain an addition temporal channel. A majority of existing works divide a video into short segments, predict relationships in each segment, and merge them. Such methods cannot capture relations involving long motions. Predicting the same relationship across neighboring video segments is also inefficient. To address these issues, this work proposes a novel sliding-window scheme to simultaneously predict short-term and long-term relationships. We run windows with different kernel sizes on object tracklets to generate sub-tracklet proposals with different duration, while the computational load is similar to that in segment-based methods. To fully utilize spatial and temporal information in videos, we construct one spatial and one temporal graph and employ Graph Convloutional Network to generate contextual embedding for tracklet proposal compatibility evaluation. We only predict relationships on highlycompatible proposal pairs. Our method achieves state-ofthe-art performance on both ImageNet-VidVRD and VidOR dataset across multiple tasks. Especially for ImageNet-VidVRD, we obtain an average of 3% (R@50 from 8.07% to 11.21%) improvement under all evaluation metrics.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Enriching Local and Global Contexts for Temporal Action LocalizationZixin Zhu, Wei Tang, Le Wang, Nanning Zheng 等ICCV 2021 · 被引用 134 次
- Video Visual Relation Detection via Iterative InferenceXindi Shang, Yicong Li, Junbin Xiao, Wei Ji 等ACM MM 2021 · 被引用 41 次
- LIGHTEN: Learning Interactions with Graph and Hierarchical TEmporal Networks for HOI in videosSai Praneeth Reddy Sunkesula, Rishabh Dabral, Ganesh RamakrishnanACM MM 2020 · 被引用 36 次
- VRDFormer: End-to-End Video Visual Relation Detection with TransformersSipeng Zheng, Shizhe Chen, Qin JinCVPR 2022 · 被引用 16 次
- SportsHHI: A Dataset for Human-Human Interaction Detection in Sports VideosTao Wu, Runyu He, Gangshan Wu, Limin WangCVPR 2024 · 被引用 9 次
相关 Paper
- Video Relation Detection via Multiple Hypothesis AssociationZixuan Su, Xindi Shang, Jingjing Chen, Yu-Gang Jiang 等ACM MM 2020 · 被引用 37 次
- Interventional Video Relation DetectionYicong Li, Xun Yang, Xindi Shang, Tat-Seng ChuaACM MM 2021 · 被引用 61 次
- VrdONE: One-stage Video Visual Relation DetectionXinjie Jiang, Chenxi Zheng, Xuemiao Xu, Bangzhen Liu 等ACM MM 2024 · 被引用 1 次
- Exploiting Better Feature Aggregation for Video Object DetectionLiang Han, Pichao Wang, Zhaozheng Yin, Fan Wang 等ACM MM 2020 · 被引用 37 次
- Visual Relation of Interest DetectionFan Yu, Haonan Wang, Tongwei Ren, Jinhui Tang 等ACM MM 2020 · 被引用 11 次
