Prototypical Action Reasoning Facilitated by Vision-Language Alignment for Egocentric Action Anticipation
Jiang Shao, Xinbo Zhao, Wenyin Tuo, Xiaochun Zou
Abstract
Egocentric Action Anticipation aims to infer future actions from videos, which is crucial for embodied AI systems. However, its advancement is hindered by the inherent stochasticity of the future, which introduces significant prediction uncertainty. Prevailing methods typically adopt an end-to-end approach to model holistic spatiotemporal contexts, yet they often lack explicit semantic reasoning capabilities, making it difficult to handle open-ended future uncertainties. To address these challenges, we propose a Prototypical Action Reasoning Framework Facilitated by Vision-Language Alignment (PAR-VLA), which leverages the semantic alignment capability of vision-language models to learn disentangled visual prototypes for verbs and nouns. These prototypes serve as robust semantic anchors, transforming the unconstrained temporal prediction problem into a conditional forecasting task guided by welldefined semantic concepts. Our multi-stage framework first extracts visually-grounded and text-aligned prototype groups from a VLM, learning multiple prototypes per category to capture intra-class diversity. Subsequently, a novel Prototypical Context Reasoning-guided Verb-Noun Encoding branch dynamically retrieves the most relevant verb and noun concepts based on visual observations and explicitly models their interactions to guide temporal anticipation. Furthermore, we introduce Dual-Stream Symbiotic Predictive Decoders to more finely capture the interdependencies between verbs and nouns during the prediction process. Experiments Results demonstrate that PAR-VLA achieves state-of-the-art performance and exhibits a strong capability in dealing with future uncertainty.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on15
- Are Language Models Actually Useful for Time Series Forecasting?Mingtian Tan, Mike A. Merrill, Vinayak Gupta, Tim Althoff et al.NeurIPS 2024 · 326 citations
- Anticipative Video TransformerRohit Girdhar, Kristen GraumanICCV 2021 · 270 citations
- What Would You Expect? Anticipating Egocentric Actions With Rolling-Unrolling LSTMs and Modality AttentionAntonino Furnari, Giovanni Maria FarinellaICCV 2019 · 204 citations
- From News to Forecast: Integrating Event Analysis in LLM-Based Time Series Forecasting with ReflectionXinlei Wang, Maike Feng, Jing Qiu, Jinjin Gu et al.NeurIPS 2024 · 181 citations
- MeMViT: Memory-Augmented Multiscale Vision Transformer for Efficient Long-Term Video RecognitionChao-Yuan Wu, Yanghao Li, Karttikeya Mangalam, Haoqi Fan et al.CVPR 2022 · 158 citations
Related papers
- Chain of World: World Model Thinking in Latent MotionFuxiang Yang, Donglin Di, Lulu Tang, Xuancheng Zhang et al.CVPR 2026 · 11 citations
- Intention-Guided Cognitive Reasoning for Egocentric Long-Term Action AnticipationQiaohui Chu, Haoyu Zhang, Meng Liu, Yisen Feng et al.AAAI 2026 · 3 citations
- De-biased Natural Language Egocentric Task Verification via Prototypical Evidence LearningChong Liu, Xun Jiang, Fumin Shen, Lei Zhu et al.AAAI 2026
- EgoTV: Egocentric Task Verification from Natural Language Task DescriptionsRishi Hazra, Brian Chen, Akshara Rai, Nitin Kamra et al.ICCV 2023
- Spatially Guided Training for Vision-Language-Action ModelJinhui Ye, Fangjing Wang, Ning Gao, Junqiu Yu et al.ICLR 2026 · 6 citations
