Social Fabric: Tubelet Compositions for Video Relation Detection
Shuo Chen, Zenglin Shi, Pascal Mettes, Cees G. M. Snoek
摘要
This paper strives to classify and detect the relationship between object tubelets appearing within a video as a 〈subject-predicate-object〉 triplet. Where existing works treat object proposals or tubelets as single entities and model their relations a posteriori, we propose to classify and detect predicates for pairs of object tubelets a priori. We also propose Social Fabric: an encoding that represents a pair of object tubelets as a composition of interaction primitives. These primitives are learned over all relations, resulting in a compact representation able to localize and classify relations from the pool of co-occurring object tubelets across all timespans in a video. The encoding enables our two-stage network. In the first stage, we train Social Fabric to suggest proposals that are likely interacting. We use the Social Fabric in the second stage to simultaneously finetune and predict predicate labels for the tubelets. Experiments demonstrate the benefit of early video relation modeling, our encoding and the two-stage architecture, leading to a new state-of-the-art on two benchmarks. We also show how the encoding enables query-by-primitive-example to search for spatio-temporal video relations. Code: https://github.com/shanshuo/Social-Fabric.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Classification-Then-Grounding: Reformulating Video Scene Graphs as Temporal Bipartite GraphsKaifeng Gao, Long Chen, Yulei Niu, Jian Shao 等CVPR 2022 · 被引用 34 次
- Tubelet-Contrastive Self-Supervision for Video-Efficient GeneralizationFida Mohammad Thoker, Hazel Doughty, Cees G. M. SnoekICCV 2023 · 被引用 13 次
- Compositional Prompt Tuning with Motion Cues for Open-vocabulary Video Relation DetectionKaifeng Gao, Long Chen, Hanwang Zhang, Jun Xiao 等ICLR 2023 · 被引用 9 次
- Triple Correlations-Guided Label Supplementation for Unbiased Video Scene Graph GenerationWenqing Wang, Kaifeng Gao, Yawei Luo, Tao Jiang 等ACM MM 2023 · 被引用 7 次
- VrdONE: One-stage Video Visual Relation DetectionXinjie Jiang, Chenxi Zheng, Xuemiao Xu, Bangzhen Liu 等ACM MM 2024 · 被引用 1 次
它引用的顶会 Paper11
- Pose-Aware Multi-Level Feature Network for Human Object Interaction DetectionBo Wan, Desen Zhou, Yongfei Liu, Rongjie Li 等ICCV 2019 · 被引用 224 次
- Video Relation Detection via Multiple Hypothesis AssociationZixuan Su, Xindi Shang, Jingjing Chen, Yu-Gang Jiang 等ACM MM 2020 · 被引用 37 次
- LIGHTEN: Learning Interactions with Graph and Hierarchical TEmporal Networks for HOI in videosSai Praneeth Reddy Sunkesula, Rishabh Dabral, Ganesh RamakrishnanACM MM 2020 · 被引用 36 次
- Learning Interactions and Relationships Between Movie CharactersAnna Kukleva, Makarand Tapaswi, Ivan LaptevCVPR 2020
- PPDM: Parallel Point Detection and Matching for Real-Time Human-Object Interaction DetectionYue Liao, Si Liu, Fei Wang, Yanjie Chen 等CVPR 2020
相关 Paper
- Localize, Assemble, and Predicate: Contextual Object Proposal Embedding for Visual Relation DetectionRuihai Wu, Kehan Xu, Chenchen Liu, Nan Zhuang 等AAAI 2020 · 被引用 7 次
- Spatio-Temporal Graph for Video Captioning With Knowledge DistillationBoxiao Pan, Haoye Cai, De-An Huang, Kuan-Hui Lee 等CVPR 2020
- Video-based Human-Object Interaction Detection from Tubelet TokensDanyang Tu, Wei Sun, Xiongkuo Min, Guangtao Zhai 等NeurIPS 2022 · 被引用 24 次
- TubeR: Tubelet Transformer for Video Action DetectionJiaojiao Zhao, Yanyi Zhang, Xinyu Li, Hao Chen 等CVPR 2022 · 被引用 77 次
- Interventional Video Relation DetectionYicong Li, Xun Yang, Xindi Shang, Tat-Seng ChuaACM MM 2021 · 被引用 61 次
