Anticipating Future Relations via Graph Growing for Action Prediction
Xinxiao Wu, Jianwei Zhao, Ruiqi Wang
摘要
Predicting actions from partially observed videos is challenging as the partial videos containing incomplete action executions have insufficient discriminative information for classification. Recent progress has been made through enriching the features of the observed video part or generating the features for the unobserved video part, but without explicitly modeling the fine-grained evolution of visual object relations over both space and time. In this paper, we investigate how the interaction and correlation between visual objects evolve and propose a graph growing method to anticipate future object relations from limited video observations for reliable action prediction. There are two tasks in our method. First, we work with spatial-temporal graph neural networks to reason object relations in the observed video part. Then, we synthesize the spatial-temporal relation representation for the unobserved video part via graph node generation and aggregation. These two tasks are jointly learned to enable the anticipated future relation representation informative to action prediction. Experimental results on two action video datasets demonstrate the effectiveness of our method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- EAST: Early Action Prediction Sampling Strategy with Token MaskingIva Sović, Ivan Martinović, Marin OršićICLR 2026
- The Wisdom of Crowds: Temporal Progressive Attention for Early Action PredictionAlexandros Stergiou, Dima DamenCVPR 2023
它引用的顶会 Paper3
- Predicting the Future: A Jointly Learnt Model for Action AnticipationHarshala Gammulle, Simon Denman, Sridha Sridharan, Clinton FookesICCV 2019 · 被引用 93 次
- Spatiotemporal Feature Residual Propagation for Action PredictionHe Zhao, Rick WildesICCV 2019 · 被引用 40 次
- Reasoning About Human-Object Interactions Through Dual Attention NetworksTete Xiao, Quanfu Fan, Danny Gutfreund, Mathew Monfort 等ICCV 2019 · 被引用 36 次
相关 Paper
- Dynamic Scene Graph Generation via Anticipatory Pre-trainingYiming Li, Xiaoshan Yang, Changsheng XuCVPR 2022 · 被引用 38 次
- Discovering Dynamic Salient Regions for Spatio-Temporal Graph Neural NetworksIulia Duta, Andrei Liviu Nicolicioiu, Marius LeordeanuNeurIPS 2021 · 被引用 8 次
- Towards Accurate 3D Human Motion Prediction From Incomplete ObservationsQiongjie Cui, Huaijiang SunCVPR 2021
- Multimodal Global Relation Knowledge Distillation for Egocentric Action AnticipationYi Huang, Xiaoshan Yang, Changsheng XuACM MM 2021 · 被引用 11 次
- Unified Graph Structured Models for Video UnderstandingAnurag Arnab, Chen Sun, Cordelia SchmidICCV 2021 · 被引用 57 次
