Anticipative Video Transformer
Rohit Girdhar, Kristen Grauman
摘要
We propose Anticipative Video Transformer (AVT), an end-to-end attention-based video modeling architecture that attends to the previously observed video in order to anticipate future actions. We train the model jointly to predict the next action in a video sequence, while also learning frame feature encoders that are predictive of successive future frames’ features. Compared to existing temporal aggregation strategies, AVT has the advantage of both maintaining the sequential progression of observed actions while still capturing long-range dependencies—both critical for the anticipation task. Through extensive experiments, we show that AVT obtains the best reported performance on four popular action anticipation benchmarks: EpicKitchens-55, EpicKitchens-100, EGTEA Gaze+, and 50-Salads; and it wins first place in the EpicKitchens-100 CVPR’21 challenge.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper89
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis 等CVPR 2022 · 被引用 525 次
- Omnivore: A Single Model for Many Visual ModalitiesRohit Girdhar, Mannat Singh, Nikhila Ravi, Laurens van der Maaten 等CVPR 2022 · 被引用 185 次
- MeMViT: Memory-Augmented Multiscale Vision Transformer for Efficient Long-Term Video RecognitionChao-Yuan Wu, Yanghao Li, Karttikeya Mangalam, Haoqi Fan 等CVPR 2022 · 被引用 158 次
- AntGPT: Can Large Language Models Help Long-term Action Anticipation from Videos?Qi Zhao, Shijie Wang, Ce Zhang, Changcheng Fu 等ICLR 2024 · 被引用 93 次
- VideoLLM-MoD: Efficient Video-Language Streaming with Mixture-of-Depths Vision ComputationShiwei Wu, Joya Chen, Kevin Qinghong Lin, Qimeng Wang 等NeurIPS 2024 · 被引用 78 次
它引用的顶会 Paper16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
相关 Paper
- Future Transformer for Long-term Action AnticipationDayoung Gong, Joonseok Lee, Manjin Kim, Seong Jong Ha 等CVPR 2022 · 被引用 56 次
- Action-Guided Attention for Video Action AnticipationTsung-Ming Tai, Sofia Casarin, Andrea Pilzer, Werner Nutt 等ICLR 2026 · 被引用 5 次
- What Would You Expect? Anticipating Egocentric Actions With Rolling-Unrolling LSTMs and Modality AttentionAntonino Furnari, Giovanni Maria FarinellaICCV 2019 · 被引用 204 次
- ActFusion: a Unified Diffusion Model for Action Segmentation and AnticipationDayoung Gong, Suha Kwak, Minsu ChoNeurIPS 2024 · 被引用 14 次
- Uncertainty-aware Action Decoupling Transformer for Action AnticipationHongji Guo, Nakul Agarwal, Shao-Yuan Lo, Kwonjoon Lee 等CVPR 2024
