Learning to Anticipate Future with Dynamic Context Removal
Xinyu Xu, Yong-Lu Li, Cewu Lu
摘要
Anticipating future events is an essential feature for in-telligent systems and embodied AI. However, compared to the traditional recognition task, the uncertainty of future and reasoning ability requirement make the anticipation task very challenging and far beyond solved. In this filed, previous methods usually care more about the model ar-chitecture design or but few attention has been put on how to train an anticipation model with a proper learning policy. To this end, in this work, we propose a novel training scheme called Dynamic Context Removal (DCR), which dynamically schedule the visibility of observed future in the learning procedure. It follows the human-like curriculum learning process, i.e., gradually removing the event context to increase the anticipation difficulty till satisfying the final anticipation target. Our learning scheme is plug-and-play and easy to integrate any reasoning model including transformer and LSTM, with advantages in both effectiveness and efficiency. In extensive experiments, the pro-posed method achieves state-of-the-art on four widely-used benchmarks. Our code and models are publicly released at https://github.com/AllenXuuuIDCR.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Memory-and-Anticipation Transformer for Online Action UnderstandingJiahao Wang, Guo Chen, Yifei Huang, Limin Wang 等ICCV 2023 · 被引用 72 次
- Beyond Object Recognition: A New Benchmark towards Object Concept LearningYonglu Li, Yue Xu, Xinyu Xu, Xiaohan Mao 等ICCV 2023 · 被引用 12 次
- GePSAn: Generative Procedure Step Anticipation in Cooking VideosMohamed Ashraf Abdelsalam, Samrudhdhi B. Rangrej, Isma Hadji, Nikita Dvornik 等ICCV 2023 · 被引用 10 次
- Towards Understanding Future: Consistency Guided Probabilistic Modeling for Action AnticipationZhao Xie, Yadong Shi, Kewei Wu, Yaru Cheng 等AAAI 2024 · 被引用 9 次
- Interacted Object Grounding in Spatio-Temporal Human-Object InteractionsXiaoyang Liu, Boran Wen, Xinpeng Liu, Zizheng Zhou 等AAAI 2025 · 被引用 6 次
它引用的顶会 Paper8
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- Multiscale Vision TransformersHaoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li 等ICCV 2021 · 被引用 1,611 次
- Video Classification With Channel-Separated Convolutional NetworksDu Tran, Heng Wang, Matt Feiszli, Lorenzo TorresaniICCV 2019 · 被引用 647 次
相关 Paper
- A-CAP: Anticipation Captioning with Commonsense KnowledgeDuc Minh Vo, Quoc-An Luong, Akihiro Sugimoto, Hideki NakayamaCVPR 2023
- From Perception to Planning: Evolving Ego-Centric Task-Oriented Spatiotemporal Reasoning via Curriculum LearningXiaoda Yang, Yuxiang Liu, Shenzhou Gao, Can Wang 等ICML 2026
- Beyond Examples: Towards Automated Thought-level In-Context Reasoning for Large Language ModelsJinyang Wu, Mingkuan Feng, Shuai Zhang, Feihu Che 等ACL 2026
- Uncertainty-aware Action Decoupling Transformer for Action AnticipationHongji Guo, Nakul Agarwal, Shao-Yuan Lo, Kwonjoon Lee 等CVPR 2024
- EDCO: Dynamic Curriculum Orchestration for Domain-specific Large Language Model Fine-tuningJing-Cheng Pang, Sun Liu, Chang Zhou, Xian Tang 等ICML 2026
