Learning to Anticipate Future with Dynamic Context Removal
Xinyu Xu, Yong-Lu Li, Cewu Lu
Abstract
Anticipating future events is an essential feature for in-telligent systems and embodied AI. However, compared to the traditional recognition task, the uncertainty of future and reasoning ability requirement make the anticipation task very challenging and far beyond solved. In this filed, previous methods usually care more about the model ar-chitecture design or but few attention has been put on how to train an anticipation model with a proper learning policy. To this end, in this work, we propose a novel training scheme called Dynamic Context Removal (DCR), which dynamically schedule the visibility of observed future in the learning procedure. It follows the human-like curriculum learning process, i.e., gradually removing the event context to increase the anticipation difficulty till satisfying the final anticipation target. Our learning scheme is plug-and-play and easy to integrate any reasoning model including transformer and LSTM, with advantages in both effectiveness and efficiency. In extensive experiments, the pro-posed method achieves state-of-the-art on four widely-used benchmarks. Our code and models are publicly released at https://github.com/AllenXuuuIDCR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Memory-and-Anticipation Transformer for Online Action UnderstandingJiahao Wang, Guo Chen, Yifei Huang, Limin Wang et al.ICCV 2023 · 72 citations
- Beyond Object Recognition: A New Benchmark towards Object Concept LearningYonglu Li, Yue Xu, Xinyu Xu, Xiaohan Mao et al.ICCV 2023 · 12 citations
- GePSAn: Generative Procedure Step Anticipation in Cooking VideosMohamed Ashraf Abdelsalam, Samrudhdhi B. Rangrej, Isma Hadji, Nikita Dvornik et al.ICCV 2023 · 10 citations
- Towards Understanding Future: Consistency Guided Probabilistic Modeling for Action AnticipationZhao Xie, Yadong Shi, Kewei Wu, Yaru Cheng et al.AAAI 2024 · 9 citations
- Interacted Object Grounding in Spatio-Temporal Human-Object InteractionsXiaoyang Liu, Boran Wen, Xinpeng Liu, Zizheng Zhou et al.AAAI 2025 · 6 citations
Builds on8
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Multiscale Vision TransformersHaoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li et al.ICCV 2021 · 1,611 citations
- Video Classification With Channel-Separated Convolutional NetworksDu Tran, Heng Wang, Matt Feiszli, Lorenzo TorresaniICCV 2019 · 647 citations
Related papers
- A-CAP: Anticipation Captioning with Commonsense KnowledgeDuc Minh Vo, Quoc-An Luong, Akihiro Sugimoto, Hideki NakayamaCVPR 2023
- From Perception to Planning: Evolving Ego-Centric Task-Oriented Spatiotemporal Reasoning via Curriculum LearningXiaoda Yang, Yuxiang Liu, Shenzhou Gao, Can Wang et al.ICML 2026
- Beyond Examples: Towards Automated Thought-level In-Context Reasoning for Large Language ModelsJinyang Wu, Mingkuan Feng, Shuai Zhang, Feihu Che et al.ACL 2026
- Uncertainty-aware Action Decoupling Transformer for Action AnticipationHongji Guo, Nakul Agarwal, Shao-Yuan Lo, Kwonjoon Lee et al.CVPR 2024
- EDCO: Dynamic Curriculum Orchestration for Domain-specific Large Language Model Fine-tuningJing-Cheng Pang, Sun Liu, Chang Zhou, Xian Tang et al.ICML 2026
