Spatiotemporal Feature Residual Propagation for Action Prediction
He Zhao, Rick Wildes
Abstract
Recognizing actions from limited preliminary video observations has seen considerable recent progress. Typically, however, such progress has been had without explicitly modeling fine-grained motion evolution as a potentially valuable information source. In this study, we address this task by investigating how action patterns evolve over time in a spatial feature space. There are three key components to our system. First, we work with intermediate-layer ConvNet features, which allow for abstraction from raw data, while retaining spatial layout, which is sacrificed in approaches that rely on vectorized global representations. Second, instead of propagating features per se, we propagate their residuals across time, which allows for a compact representation that reduces redundancy while retaining essential information about evolution over time. Third, we employ a Kalman filter to combat error build-up and unify across prediction start times. Extensive experimental results on the JHMDB21, UCF101 and BIT datasets show that our approach leads to a new state-of-the-art in action prediction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Bifold and Semantic Reasoning for Pedestrian Behavior PredictionAmir Rasouli, Mohsen Rohani, Jun LuoICCV 2021 · 69 citations
- Distillation Using Oracle Queries for Transformer-based Human-Object Interaction DetectionXian Qu, Changxing Ding, Xingao Li, Xubin Zhong et al.CVPR 2022 · 48 citations
- Anticipating Future Relations via Graph Growing for Action PredictionXinxiao Wu, Jianwei Zhao, Ruiqi WangAAAI 2021 · 24 citations
- EAST: Early Action Prediction Sampling Strategy with Token MaskingIva Sović, Ivan Martinović, Marin OršićICLR 2026
- Anticipating Human Actions by Correlating Past With the Future With Jaccard Similarity MeasuresBasura Fernando, Samitha HerathCVPR 2021
Related papers
- STM: SpatioTemporal and Motion Encoding for Action RecognitionBoyuan Jiang, Mengmeng Wang, Weihao Gan, Wei Wu et al.ICCV 2019 · 442 citations
- TEA: Temporal Excitation and Aggregation for Action RecognitionYan Li, Bin Ji, Xintian Shi, Jianguo Zhang et al.CVPR 2020
- The Wisdom of Crowds: Temporal Progressive Attention for Early Action PredictionAlexandros Stergiou, Dima DamenCVPR 2023
- Evolving Space-Time Neural Architectures for VideosA. J. Piergiovanni, Anelia Angelova, Alexander Toshev, Michael S. RyooICCV 2019 · 62 citations
- V4D: 4D Convolutional Neural Networks for Video-level Representation LearningShiwen Zhang, Sheng Guo, Weilin Huang, Matthew R. Scott et al.ICLR 2020 · 81 citations
