Detecting Human-Object Relationships in Videos
Jingwei Ji, Rishi Desai, Juan Carlos Niebles
Abstract
We study a crucial problem in video analysis: human-object relationship detection. The majority of previous approaches are developed only for the static image scenario, without incorporating the temporal dynamics so vital to contextualizing human-object relationships. We propose a model with Intra- and Inter-Transformers, enabling joint spatial and temporal reasoning on multiple visual concepts of objects, relationships, and human poses. We find that applying attention mechanisms among features distributed spatio-temporally greatly improves our understanding of human-object relationships. Our method is validated on two datasets, Action Genome and CAD-120-EVAR, and achieves state-of-the-art performance on both of them.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c6130cd3-0ae6-418a-9c03-9a459dc1c4ffCited by top-tier papers10
- InterDiff: Generating 3D Human-Object Interactions with Physics-Informed DiffusionSirui Xu, Zhengyuan Li, Yu-Xiong Wang, Liang-Yan GuiICCV 2023 · 201 citations
- Video-based Human-Object Interaction Detection from Tubelet TokensDanyang Tu, Wei Sun, Xiongkuo Min, Guangtao Zhai et al.NeurIPS 2022 · 24 citations
- OED: Towards One-stage End-to-End Dynamic Scene Graph GenerationGuan Wang, Zhimin Li, Qingchao Chen, Yang LiuCVPR 2024 · 12 citations
- SportsHHI: A Dataset for Human-Human Interaction Detection in Sports VideosTao Wu, Runyu He, Gangshan Wu, Limin WangCVPR 2024 · 9 citations
- TRKT: Weakly Supervised Dynamic Scene Graph Generation with Temporal-Enhanced Relation-Aware Knowledge TransferringZhu Xu, Ting Lei, Zhimin Li, Guan Wang et al.ICCV 2025 · 3 citations
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy et al.ICCV 2019 · 1,396 citations
- Graph Transformer for Graph-to-Sequence LearningDeng Cai, Wai LamAAAI 2020 · 247 citations
- Pose-Aware Multi-Level Feature Network for Human Object Interaction DetectionBo Wan, Desen Zhou, Yongfei Liu, Rongjie Li et al.ICCV 2019 · 224 citations
Related papers
- Dynamic Scene Graph Generation via Anticipatory Pre-trainingYiming Li, Xiaoshan Yang, Changsheng XuCVPR 2022 · 38 citations
- Unified Graph Structured Models for Video UnderstandingAnurag Arnab, Chen Sun, Cordelia SchmidICCV 2021 · 57 citations
- Spatial-Temporal Transformer for Dynamic Scene Graph GenerationYuren Cong, Wentong Liao, Hanno Ackermann, Bodo Rosenhahn et al.ICCV 2021 · 163 citations
- What to look at and where: Semantic and Spatial Refined Transformer for detecting human-object interactionsA. S. M. Iftekhar, Hao Chen, Kaustav Kundu, Xinyu Li et al.CVPR 2022 · 50 citations
- VRDFormer: End-to-End Video Visual Relation Detection with TransformersSipeng Zheng, Shizhe Chen, Qin JinCVPR 2022 · 16 citations
