The Wisdom of Crowds: Temporal Progressive Attention for Early Action Prediction
Alexandros Stergiou, Dima Damen
Abstract
Early action prediction deals with inferring the ongoing action from partially-observed videos, typically at the outset of the video. We propose a bottleneck-based attention model that captures the evolution of the action, through progressive sampling over fine-to-coarse scales. Our proposed Temporal Progressive (TemPr) model is composed of multiple attention towers, one for each scale. The predicted action label is based on the collective agreement considering confidences of these towers. Extensive experiments over four video datasets showcase state-of-the-art performance on the task of Early Action Prediction across a range of encoder architectures. We demonstrate the effectiveness and consistency of TemPr through detailed ablations. †
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- GenRec: Unifying Video Generation and Recognition with Diffusion ModelsZejia Weng, Xitong Yang, Zhen Xing, Zuxuan Wu et al.NeurIPS 2024 · 19 citations
- Fostering Video Reasoning via Next-Event PredictionHaonan Wang, Hongfu Liu, Xiangyan Liu, Chao Du et al.ICLR 2026 · 14 citations
- EAST: Early Action Prediction Sampling Strategy with Token MaskingIva Sović, Ivan Martinović, Marin OršićICLR 2026
Builds on20
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei et al.CVPR 2022 · 1,847 citations
Related papers
- End-to-End Spatio-Temporal Action Localisation with Video TransformersAlexey A. Gritsenko, Xuehan Xiong, Josip Djolonga, Mostafa Dehghani et al.CVPR 2024
- Future Transformer for Long-term Action AnticipationDayoung Gong, Joonseok Lee, Manjin Kim, Seong Jong Ha et al.CVPR 2022 · 56 citations
- How Much Temporal Long-Term Context is Needed for Action Segmentation?Emad Bahrami Rad, Gianpiero Francesca, Juergen GallICCV 2023 · 54 citations
- Progressive Boundary Refinement Network for Temporal Action DetectionQinying Liu, Zilei WangAAAI 2020 · 156 citations
- Towards Understanding Future: Consistency Guided Probabilistic Modeling for Action AnticipationZhao Xie, Yadong Shi, Kewei Wu, Yaru Cheng et al.AAAI 2024 · 9 citations
