SEPT: Towards Efficient Scene Representation Learning for Motion Prediction
Zhiqian Lan, Yuxuan Jiang, Yao Mu, Chen Chen, Shengbo Eben Li
Abstract
Motion prediction is crucial for autonomous vehicles to operate safely in complex traffic environments. Extracting effective spatiotemporal relationships among traffic elements is key to accurate forecasting. Inspired by the successful practice of pretrained large language models, this paper presents SEPT, a modeling framework that leverages self-supervised learning to develop powerful spatiotemporal understanding for complex traffic scenes. Specifically, our approach involves three masking-reconstruction modeling tasks on scene inputs including agents' trajectories and road network, pretraining the scene encoder to capture kinematics within trajectory, spatial structure of road network, and interactions among roads and agents. The pretrained encoder is then finetuned on the downstream forecasting task. Extensive experiments demonstrate that SEPT, without elaborate architectural design or manual feature engineering, achieves state-of-the-art performance on the Argoverse 1 and Argoverse 2 motion forecasting benchmarks, outperforming previous methods on all main metrics by a large margin.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 73fdc76e-4fe7-4e33-8ad1-718e678a5b04Cited by top-tier papers12
- DeMo: Decoupling Motion Forecasting into Directional Intentions and Dynamic StatesBozhou Zhang, Nan Song, Li ZhangNeurIPS 2024 · 34 citations
- Motion Forecasting in Continuous DrivingNan Song, Bozhou Zhang, Xiatian Zhu, Li ZhangNeurIPS 2024 · 33 citations
- REDOUBT: Duo Safety Validation for Autonomous Vehicle Motion PlanningShuguang Wang, Qian Zhou, Kui Wu, Dapeng Wu et al.NeurIPS 2025 · 6 citations
- Recover to Predict: Progressive Retrospective Learning for Variable-Length Trajectory PredictionHao Zhou, Lu Qi, Xiangtai Li, Jie Zhang et al.CVPR 2026 · 2 citations
- UniMotion: A Unified Motion Framework for Simulation, Prediction and PlanningNan Song, Junzhe Jiang, Jingyu Li, Xiatian Zhu et al.NeurIPS 2025 · 2 citations
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- DenseTNT: End-to-end Trajectory Prediction from Dense Goal SetsJunru Gu, Chen Sun, Hang ZhaoICCV 2021 · 563 citations
- Motion Transformer with Global Intention Localization and Local Movement RefinementShaoshuai Shi, Li Jiang, Dengxin Dai, Bernt SchieleNeurIPS 2022 · 515 citations
- HiVT: Hierarchical Vector Transformer for Multi-Agent Motion PredictionZikang Zhou, Luyao Ye, Jianping Wang, Kui Wu et al.CVPR 2022 · 379 citations
- Scene Transformer: A unified architecture for predicting future trajectories of multiple agentsJiquan Ngiam, Vijay Vasudevan, Benjamin Caine, Zhengdong Zhang et al.ICLR 2022 · 194 citations
Related papers
- Forecast-MAE: Self-supervised Pre-training for Motion Forecasting with Masked AutoencodersJie Cheng, Xiaodong Mei, Ming LiuICCV 2023 · 123 citations
- Query-Centric Trajectory PredictionZikang Zhou, Jianping Wang, Yung-Hui Li, Yu-Kai HuangCVPR 2023
- DONUT: A Decoder-Only Model for Trajectory PredictionMarkus Knoche, Daan de Geus, Bastian LeibeICCV 2025
- LaPred: Lane-Aware Prediction of Multi-Modal Future Trajectories of Dynamic AgentsByeoungdo Kim, SeongHyeon Park, Seokhwan Lee, Elbek Khoshimjonov et al.CVPR 2021
- Future-Aware Interaction Network for Motion ForecastingShijie Li, Chunyu Liu, Xun Xu, Si Yong Yeo et al.ICCV 2025 · 4 citations
