LIGHTEN: Learning Interactions with Graph and Hierarchical TEmporal Networks for HOI in videos
Sai Praneeth Reddy Sunkesula, Rishabh Dabral, Ganesh Ramakrishnan
Abstract
Analyzing the interactions between humans and objects from a video includes identification of the relationships between humans and the objects present in the video. It can be thought of as a specialized version of Visual Relationship Detection, wherein one of the objects must be a human. While traditional methods formulate the problem as inference on a sequence of video segments, we present a hierarchical approach, LIGHTEN, to learn visual features to effectively capture spatio-temporal cues at multiple granularities in a video. Unlike current approaches, LIGHTEN avoids using ground truth data like depth maps or 3D human pose, thus increasing generalization across non-RGBD datasets as well. Furthermore, we achieve the same using only the visual features, instead of the commonly used hand-crafted spatial features. We achieve state-of-the-art results in human-object interaction detection (88.9% and 92.6%) and anticipation tasks of CAD-120 and competitive results on image based HOI detection in V-COCO dataset, setting a new benchmark for visual features based approaches. Code for LIGHTEN is available at https://github.com/praneeth11009/LIGHTEN-Learning-Interactions-with-Graphs-and-Hierarchical-TEmporal-Networks-for-HOI
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e176ab47-0210-456f-bed8-fd2f7ea49fbcCited by top-tier papers5
- Spatio-Temporal Interaction Graph Parsing Networks for Human-Object Interaction RecognitionNing Wang, Guangming Zhu, Liang Zhang, Peiyi Shen et al.ACM MM 2021 · 32 citations
- Social Fabric: Tubelet Compositions for Video Relation DetectionShuo Chen, Zenglin Shi, Pascal Mettes, Cees G. M. SnoekICCV 2021 · 25 citations
- Video-based Human-Object Interaction Detection from Tubelet TokensDanyang Tu, Wei Sun, Xiongkuo Min, Guangtao Zhai et al.NeurIPS 2022 · 24 citations
- Open Set Video HOI detection from Action-centric Chain-of-Look PromptingNan Xi, Jingjing Meng, Junsong YuanICCV 2023 · 10 citations
- Person in Place: Generating Associative Skeleton-Guidance Maps for Human-Object Interaction Image EditingChangHee Yang, Chanhee Kang, Kyeongbo Kong, Hanni Oh et al.CVPR 2024
Builds on9
- XNect: real-time multi-person 3D motion capture with a single RGB cameraDushyant Mehta, Oleksandr Sotnychenko, Franziska Mueller, Weipeng Xu et al.SIGGRAPH 2020 · 267 citations
- Pose-Aware Multi-Level Feature Network for Human Object Interaction DetectionBo Wan, Desen Zhou, Yongfei Liu, Rongjie Li et al.ICCV 2019 · 224 citations
- Robust Data Programming with Precision-guided Labeling FunctionsOishik Chatterjee, Ganesh Ramakrishnan, Sunita SarawagiAAAI 2020 · 20 citations
- EmotiCon: Context-Aware Multimodal Emotion Recognition Using Frege's PrincipleTrisha Mittal, Pooja Guhan, Uttaran Bhattacharya, Rohan Chandra et al.CVPR 2020
- Semantics-Guided Neural Networks for Efficient Skeleton-Based Human Action RecognitionPengfei Zhang, Cuiling Lan, Wenjun Zeng, Junliang Xing et al.CVPR 2020
Related papers
- Deep Contextual Attention for Human-Object Interaction DetectionTiancai Wang, Rao Muhammad Anwer, Muhammad Haris Khan, Fahad Shahbaz Khan et al.ICCV 2019 · 130 citations
- Learning Human-Object Interaction Detection Using Interaction PointsTiancai Wang, Tong Yang, Martin Danelljan, Fahad Shahbaz Khan et al.CVPR 2020
- HORP: Human-Object Relation Priors Guided HOI DetectionPei Geng, Jian Yang, Shanshan ZhangCVPR 2025
- Exploiting Scene Graphs for Human-Object Interaction DetectionTao He, Lianli Gao, Jingkuan Song, Yuan-Fang LiICCV 2021 · 40 citations
- Detecting Human-Object Relationships in VideosJingwei Ji, Rishi Desai, Juan Carlos NieblesICCV 2021 · 47 citations
