An Egocentric Action Anticipation Framework via Fusing Intuition and Analysis
Tianyu Zhang, Weiqing Min, Ying Zhu, Yong Rui, Shuqiang Jiang
Abstract
In this paper, we focus on egocentric action anticipation from videos, which enables various applications, such as helping intelligent wearable assistants understand users' needs and enhance their capabilities in the interaction process. It requires intelligent systems to observe from the perspective of the first person and predict an action before it occurs. Owing to the uncertainty of future, it is insufficient to perform action anticipation relying on visual information especially when there exists salient visual difference between past and future. In order to alleviate this problem, which we call visual gap in this paper, we propose one novel Intuition-Analysis Integrated (IAI) framework inspired by psychological research, which mainly consists of three parts: Intuition-based Prediction Network (IPN), Analysis-based Prediction Network (APN) and Adaptive Fusion Network (AFN). To imitate the implicit intuitive thinking process, we model IPN as an encoder-decoder structure and introduce one procedural instruction learning strategy implemented by textual pre-training. On the other hand, we allow APN to process information under designed rules to imitate the explicit analytical thinking, which is divided into three steps: recognition, transitions and combination. Both the procedural instruction learning strategy in IPN and the transition step of APN are crucial to improving the anticipation performance via mitigating the visual gap problem. Considering the complementarity of intuition and analysis, AFN adopts attention fusion to adaptively integrate predictions from IPN and APN to produce the final anticipation results. We conduct experiments on the largest egocentric video dataset. Qualitative and quantitative evaluation results validate the effectiveness of our IAI framework, and demonstrate the advantage of bridging visual gap by utilizing multi-modal information, including both visual features of observed segments and sequential instructions of actions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Egocentric Prediction of Action Target in 3DYiming Li, Ziang Cao, Andrew Liang, Benjamin Liang et al.CVPR 2022 · 20 citations
- Anticipating Human Actions by Correlating Past With the Future With Jaccard Similarity MeasuresBasura Fernando, Samitha HerathCVPR 2021
Builds on1
Related papers
- Intention-Guided Cognitive Reasoning for Egocentric Long-Term Action AnticipationQiaohui Chu, Haoyu Zhang, Meng Liu, Yisen Feng et al.AAAI 2026 · 3 citations
- Summarize the Past to Predict the Future: Natural Language Descriptions of Context Boost Multimodal Object Interaction AnticipationRazvan-George Pasca, Alexey Gavryushin, Muhammad Hamza, Yen-Ling Kuo et al.CVPR 2024 · 9 citations
- Multimodal Global Relation Knowledge Distillation for Egocentric Action AnticipationYi Huang, Xiaoshan Yang, Changsheng XuACM MM 2021 · 11 citations
- Event-Guided Procedure Planning from Instructional Videos with Text SupervisionAn-Lan Wang, Kun-Yu Lin, Jia-Run Du, Jingke Meng et al.ICCV 2023 · 21 citations
- Exo2Ego: Exocentric Knowledge Guided MLLM for Egocentric Video UnderstandingHaoyu Zhang, Qiaohui Chu, Meng Liu, Haoxiang Shi et al.AAAI 2026 · 17 citations
