Uncertainty-aware Action Decoupling Transformer for Action Anticipation
Hongji Guo, Nakul Agarwal, Shao-Yuan Lo, Kwonjoon Lee, Qiang Ji
摘要
Human action anticipation aims at predicting what people will do in the future based on past observations. In this paper, we introduce Uncertainty-aware Action Decoupling Transformer (UADT) for action anticipation. Unlike existing methods that directly predict action in a verb-noun pair format, we decouple the action anticipation task into verb and noun anticipations separately. The objective is to make the two decoupled tasks assist each other and eventually improve the action anticipation task. Specifically, we propose a two-stream Transformer-based architecture which is composed of a verb-to-noun model and a noun-to-verb model. The verb-to-noun model leverages the verb information to improve the noun prediction and the other way around. We extend the model in a probabilistic manner and quantify the predictive uncertainty of each decoupled task to select features. In this way, the noun prediction leverages the most informative and redundancy-free verb features and verb prediction works similarly. Finally, the two streams are combined dynamically based on their uncertainties to make the joint action anticipation. We demonstrate the efficacy of our method by achieving state-of-the-art performance on action anticipation benchmarks including EPIC-KITCHENS, EGTEA Gaze+, and 50-Salads.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Can't make an Omelette without Breaking some Eggs: Plausible Action Anticipation using Large Video-Language ModelsHimangi Mittal, Nakul Agarwal, Shao-Yuan Lo, Kwonjoon LeeCVPR 2024 · 被引用 14 次
- Modeling Multiple Normal Action Representations for Error Detection in Procedural TasksWei-Jin Huang, Yuan-Ming Li, Zhi-Wei Xia, Yu-Ming Tang 等CVPR 2025
- Prototypical Action Reasoning Facilitated by Vision-Language Alignment for Egocentric Action AnticipationJiang Shao, Xinbo Zhao, Wenyin Tuo, Xiaochun ZouCVPR 2026
- Action Sequence Augmentation for Action AnticipationYihui Qiu, Deepu RajanICLR 2025
- MANTA: Diffusion Mamba for Efficient and Effective Stochastic Long-Term Dense Action AnticipationOlga Zatsarynna, Emad Bahrami, Yazan Abu Farha, Gianpiero Francesca 等CVPR 2025
它引用的顶会 Paper18
- Multiscale Vision TransformersHaoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li 等ICCV 2021 · 被引用 1,611 次
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis 等CVPR 2022 · 被引用 525 次
- Uncertainty-Guided Transformer Reasoning for Camouflaged Object DetectionFan Yang, Qiang Zhai, Xin Li, Rui Huang 等ICCV 2021 · 被引用 293 次
- Anticipative Video TransformerRohit Girdhar, Kristen GraumanICCV 2021 · 被引用 270 次
- What Would You Expect? Anticipating Egocentric Actions With Rolling-Unrolling LSTMs and Modality AttentionAntonino Furnari, Giovanni Maria FarinellaICCV 2019 · 被引用 204 次
相关 Paper
- Future Transformer for Long-term Action AnticipationDayoung Gong, Joonseok Lee, Manjin Kim, Seong Jong Ha 等CVPR 2022 · 被引用 56 次
- Uncertainty-Guided Probabilistic Transformer for Complex Action RecognitionHongji Guo, Hanjing Wang, Qiang JiCVPR 2022 · 被引用 42 次
- Memory-and-Anticipation Transformer for Online Action UnderstandingJiahao Wang, Guo Chen, Yifei Huang, Limin Wang 等ICCV 2023 · 被引用 72 次
- ActFusion: a Unified Diffusion Model for Action Segmentation and AnticipationDayoung Gong, Suha Kwak, Minsu ChoNeurIPS 2024 · 被引用 14 次
- AntGPT: Can Large Language Models Help Long-term Action Anticipation from Videos?Qi Zhao, Shijie Wang, Ce Zhang, Changcheng Fu 等ICLR 2024 · 被引用 93 次
