Latency Matters: Real-Time Action Forecasting Transformer
Harshayu Girase, Nakul Agarwal, Chiho Choi, Karttikeya Mangalam
摘要
We present RAFTformer, a real-time action forecasting transformer for latency-aware real-world action forecasting. RAFTformer is a two-stage fully transformer based architecture comprising of a video transformer backbone that operates on high resolution, short-range clips, and a head transformer encoder that temporally aggregates information from multiple short-range clips to span a long-term horizon. Additionally, we propose a novel self-supervised shuffled causal masking scheme as a model level augmentation to improve forecasting fidelity. Finally, we also propose a novel real-time evaluation setting for action forecasting that directly couples model inference latency to overall forecasting performance and brings forth a hitherto overlooked trade-off between latency and action forecasting performance. Our parsimonious network design facilitates RAFTformer inference latency to be 9× smaller than prior works at the same forecasting accuracy. Owing to its two-staged design, RAFTformer uses 94% less training compute and 90% lesser training parameters to outperform prior state-of-the-art baselines by 4.9 points on EGTEA Gaze+ and by 1.4 points on EPIC-Kitchens-100 validation set, as measured by Top-5 recall (T5R) in the offline setting. In the real-time setting, RAFTformer outperforms prior works by an even greater margin of upto 4.4 T5R points on the EPIC-Kitchens-100 dataset. Project Webpage: https://karttikeya.github . io/publication/RAFTformer/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- AntGPT: Can Large Language Models Help Long-term Action Anticipation from Videos?Qi Zhao, Shijie Wang, Ce Zhang, Changcheng Fu 等ICLR 2024 · 被引用 93 次
- Can't make an Omelette without Breaking some Eggs: Plausible Action Anticipation using Large Video-Language ModelsHimangi Mittal, Nakul Agarwal, Shao-Yuan Lo, Kwonjoon LeeCVPR 2024 · 被引用 14 次
- Summarize the Past to Predict the Future: Natural Language Descriptions of Context Boost Multimodal Object Interaction AnticipationRazvan-George Pasca, Alexey Gavryushin, Muhammad Hamza, Yen-Ling Kuo 等CVPR 2024 · 被引用 9 次
- Action-Guided Attention for Video Action AnticipationTsung-Ming Tai, Sofia Casarin, Andrea Pilzer, Werner Nutt 等ICLR 2026 · 被引用 5 次
- RiskProp: Collision-Anchored Self-Supervised Risk Propagation For Early Accident AnticipationYiyang Zou, Tianhao Zhao, Peilun Xiao, Hongyu Jin 等CVPR 2026 · 被引用 4 次
它引用的顶会 Paper19
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 被引用 2,336 次
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei 等CVPR 2022 · 被引用 1,847 次
相关 Paper
- Anticipative Video TransformerRohit Girdhar, Kristen GraumanICCV 2021 · 被引用 270 次
- Learning Streaming Video Representation via Multitask TrainingYibin Yan, Jilan Xu, Shangzhe Di, Yikun Liu 等ICCV 2025 · 被引用 1 次
- MeMViT: Memory-Augmented Multiscale Vision Transformer for Efficient Long-Term Video RecognitionChao-Yuan Wu, Yanghao Li, Karttikeya Mangalam, Haoqi Fan 等CVPR 2022 · 被引用 158 次
- Joint Hand Motion and Interaction Hotspots Prediction from Egocentric VideosShaowei Liu, Subarna Tripathi, Somdeb Majumdar, Xiaolong WangCVPR 2022 · 被引用 69 次
- Towards Lightweight Time Series Forecasting: A Patch-Wise Transformer with Weak Data EnrichingMeng Wang, Jintao Yang, Bin Yang, Hui Li 等ICDE 2025 · 被引用 10 次
