Hierarchical Video Prediction Using Relational Layouts for Human-Object Interactions
Navaneeth Bodla, Gaurav Shrivastava, Rama Chellappa, Abhinav Shrivastava
Abstract
Learning to model and predict how humans interact with objects while performing an action is challenging, and most of the existing video prediction models are ineffective in modeling complicated human-object interactions. Our work builds on hierarchical video prediction models, which disentangle the video generation process into two stages: predicting a high-level representation, such as pose sequence, and then learning a pose-to-pixels translation model for pixel generation. An action sequence for a human-object interaction task is typically very complicated, involving the evolution of pose, person's appearance, object locations, and object appearances over time. To this end, we propose a Hierarchical Video Prediction model using Relational Layouts. In the first stage, we learn to predict a sequence of layouts. A layout is a high-level representation of the video containing both pose and objects' information for every frame. The layout sequence is learned by modeling the relationships between the pose and objects using relational reasoning and recurrent neural networks. The layout sequence acts as a strong structure prior to the second stage that learns to map the layouts into pixel space. Experimental evaluation of our method on two datasets, UMD-HOI and Bimanual, shows significant improvements in standard video evaluation metrics such as LPIPS, PSNR, and SSIM. We also perform a detailed qualitative analysis of our model to demonstrate various generalizations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8a0e5f77-6537-4ceb-a5ec-39ead0004b9cCited by top-tier papers4
- Iso-Dream: Isolating and Leveraging Noncontrollable Visual Dynamics in World ModelsMinting Pan, Xiangming Zhu, Yunbo Wang, Xiaokang YangNeurIPS 2022 · 74 citations
- Video Dynamics Prior: An Internal Learning Approach for Robust Video EnhancementsGaurav Shrivastava, Ser Nam Lim, Abhinav ShrivastavaNeurIPS 2023 · 14 citations
- Video Decomposition Prior: Editing Videos Layer by LayerGaurav Shrivastava, Ser-Nam Lim, Abhinav ShrivastavaICLR 2024 · 11 citations
- Video Prediction by Modeling Videos as Continuous Multi-Dimensional ProcessesGaurav Shrivastava, Abhinav ShrivastavaCVPR 2024 · 4 citations
Related papers
- Revisiting Hierarchical Approach for Persistent Long-Term Video PredictionWonkwang Lee, Whie Jung, Han Zhang, Ting Chen et al.ICLR 2021 · 29 citations
- OneHOI: Unifying Human-Object Interaction Generation and EditingJiun Tian Hoe, Weipeng Hu, Xudong Jiang, Yap-Peng Tan et al.CVPR 2026
- Detecting Human-Object Relationships in VideosJingwei Ji, Rishi Desai, Juan Carlos NieblesICCV 2021 · 47 citations
- Learning Semantic-Aware Dynamics for Video PredictionXinzhu Bei, Yanchao Yang, Stefano SoattoCVPR 2021
- Learning from Easy to Hard Pairs: Multi-step Reasoning Network for Human-Object Interaction DetectionYuchen Zhou, Guang Tan, Mengtang Li, Chao GouACM MM 2023 · 16 citations
