Forecast-MAE: Self-supervised Pre-training for Motion Forecasting with Masked Autoencoders
Jie Cheng, Xiaodong Mei, Ming Liu
Abstract
This study explores the application of self-supervised learning (SSL) to the task of motion forecasting, an area that has not yet been extensively investigated despite the widespread success of SSL in computer vision and natural language processing. To address this gap, we introduce Forecast-MAE, an extension of the mask autoencoders framework that is specifically designed for self-supervised learning of the motion forecasting task. Our approach includes a novel masking strategy that leverages the strong interconnections between agents' trajectories and road networks, involving complementary masking of agents' future or history trajectories and random masking of lane segments. Our experiments on the challenging Argoverse 2 motion forecasting benchmark show that Forecast-MAE, which utilizes standard Transformer blocks with minimal inductive bias, achieves competitive performance compared to state-of-the-art methods that rely on supervised learning and sophisticated designs. Moreover, it outperforms the previous self-supervised learning method by a significant margin. Code is available at https://github.com/ jchengai/forecast-mae .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers29
- HPNet: Dynamic Trajectory Forecasting with Historical Prediction AttentionXiaolong Tang, Meina Kan, Shiguang Shan, Zhilong Ji et al.CVPR 2024 · 66 citations
- SEPT: Towards Efficient Scene Representation Learning for Motion PredictionZhiqian Lan, Yuxuan Jiang, Yao Mu, Chen Chen et al.ICLR 2024 · 56 citations
- DeMo: Decoupling Motion Forecasting into Directional Intentions and Dynamic StatesBozhou Zhang, Nan Song, Li ZhangNeurIPS 2024 · 34 citations
- Motion Forecasting in Continuous DrivingNan Song, Bozhou Zhang, Xiatian Zhu, Li ZhangNeurIPS 2024 · 33 citations
- DrivingRecon: Large 4D Gaussian Reconstruction Model For Autonomous DrivingHao Lu, Tianshuo Xu, Wenzhao Zheng, Yunpeng Zhang et al.NeurIPS 2025 · 26 citations
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Large Scale Interactive Motion Forecasting for Autonomous Driving : The Waymo Open Motion DatasetScott Ettinger, Shuyang Cheng, Benjamin Caine, Chenxi Liu et al.ICCV 2021 · 817 citations
Related papers
- Future-Aware Interaction Network for Motion ForecastingShijie Li, Chunyu Liu, Xun Xu, Si Yong Yeo et al.ICCV 2025 · 4 citations
- DONUT: A Decoder-Only Model for Trajectory PredictionMarkus Knoche, Daan de Geus, Bastian LeibeICCV 2025
- LTP: Lane-based Trajectory Prediction for Autonomous DrivingJingke Wang, Tengju Ye, Ziqing Gu, Junbo ChenCVPR 2022 · 75 citations
- SmartPretrain: Model-Agnostic and Dataset-Agnostic Representation Learning for Motion PredictionYang Zhou, Hao Shao, Letian Wang, Steven L. Waslander et al.ICLR 2025
- Learn TAROT with MENTOR: A Meta-Learned Self-supervised Approach for Trajectory PredictionMozhgan Pourkeshavarz, Changhe Chen, Amir RasouliICCV 2023 · 18 citations
