Prototypical Action Reasoning Facilitated by Vision-Language Alignment for Egocentric Action Anticipation
Jiang Shao, Xinbo Zhao, Wenyin Tuo, Xiaochun Zou
摘要
Egocentric Action Anticipation aims to infer future actions from videos, which is crucial for embodied AI systems. However, its advancement is hindered by the inherent stochasticity of the future, which introduces significant prediction uncertainty. Prevailing methods typically adopt an end-to-end approach to model holistic spatiotemporal contexts, yet they often lack explicit semantic reasoning capabilities, making it difficult to handle open-ended future uncertainties. To address these challenges, we propose a Prototypical Action Reasoning Framework Facilitated by Vision-Language Alignment (PAR-VLA), which leverages the semantic alignment capability of vision-language models to learn disentangled visual prototypes for verbs and nouns. These prototypes serve as robust semantic anchors, transforming the unconstrained temporal prediction problem into a conditional forecasting task guided by welldefined semantic concepts. Our multi-stage framework first extracts visually-grounded and text-aligned prototype groups from a VLM, learning multiple prototypes per category to capture intra-class diversity. Subsequently, a novel Prototypical Context Reasoning-guided Verb-Noun Encoding branch dynamically retrieves the most relevant verb and noun concepts based on visual observations and explicitly models their interactions to guide temporal anticipation. Furthermore, we introduce Dual-Stream Symbiotic Predictive Decoders to more finely capture the interdependencies between verbs and nouns during the prediction process. Experiments Results demonstrate that PAR-VLA achieves state-of-the-art performance and exhibits a strong capability in dealing with future uncertainty.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Are Language Models Actually Useful for Time Series Forecasting?Mingtian Tan, Mike A. Merrill, Vinayak Gupta, Tim Althoff 等NeurIPS 2024 · 被引用 326 次
- Anticipative Video TransformerRohit Girdhar, Kristen GraumanICCV 2021 · 被引用 270 次
- What Would You Expect? Anticipating Egocentric Actions With Rolling-Unrolling LSTMs and Modality AttentionAntonino Furnari, Giovanni Maria FarinellaICCV 2019 · 被引用 204 次
- From News to Forecast: Integrating Event Analysis in LLM-Based Time Series Forecasting with ReflectionXinlei Wang, Maike Feng, Jing Qiu, Jinjin Gu 等NeurIPS 2024 · 被引用 181 次
- MeMViT: Memory-Augmented Multiscale Vision Transformer for Efficient Long-Term Video RecognitionChao-Yuan Wu, Yanghao Li, Karttikeya Mangalam, Haoqi Fan 等CVPR 2022 · 被引用 158 次
相关 Paper
- Chain of World: World Model Thinking in Latent MotionFuxiang Yang, Donglin Di, Lulu Tang, Xuancheng Zhang 等CVPR 2026 · 被引用 11 次
- Intention-Guided Cognitive Reasoning for Egocentric Long-Term Action AnticipationQiaohui Chu, Haoyu Zhang, Meng Liu, Yisen Feng 等AAAI 2026 · 被引用 3 次
- De-biased Natural Language Egocentric Task Verification via Prototypical Evidence LearningChong Liu, Xun Jiang, Fumin Shen, Lei Zhu 等AAAI 2026
- EgoTV: Egocentric Task Verification from Natural Language Task DescriptionsRishi Hazra, Brian Chen, Akshara Rai, Nitin Kamra 等ICCV 2023
- Spatially Guided Training for Vision-Language-Action ModelJinhui Ye, Fangjing Wang, Ning Gao, Junqiu Yu 等ICLR 2026 · 被引用 6 次
