SEPT: Towards Efficient Scene Representation Learning for Motion Prediction
Zhiqian Lan, Yuxuan Jiang, Yao Mu, Chen Chen, Shengbo Eben Li
摘要
Motion prediction is crucial for autonomous vehicles to operate safely in complex traffic environments. Extracting effective spatiotemporal relationships among traffic elements is key to accurate forecasting. Inspired by the successful practice of pretrained large language models, this paper presents SEPT, a modeling framework that leverages self-supervised learning to develop powerful spatiotemporal understanding for complex traffic scenes. Specifically, our approach involves three masking-reconstruction modeling tasks on scene inputs including agents' trajectories and road network, pretraining the scene encoder to capture kinematics within trajectory, spatial structure of road network, and interactions among roads and agents. The pretrained encoder is then finetuned on the downstream forecasting task. Extensive experiments demonstrate that SEPT, without elaborate architectural design or manual feature engineering, achieves state-of-the-art performance on the Argoverse 1 and Argoverse 2 motion forecasting benchmarks, outperforming previous methods on all main metrics by a large margin.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- DeMo: Decoupling Motion Forecasting into Directional Intentions and Dynamic StatesBozhou Zhang, Nan Song, Li ZhangNeurIPS 2024 · 被引用 34 次
- Motion Forecasting in Continuous DrivingNan Song, Bozhou Zhang, Xiatian Zhu, Li ZhangNeurIPS 2024 · 被引用 33 次
- REDOUBT: Duo Safety Validation for Autonomous Vehicle Motion PlanningShuguang Wang, Qian Zhou, Kui Wu, Dapeng Wu 等NeurIPS 2025 · 被引用 6 次
- Recover to Predict: Progressive Retrospective Learning for Variable-Length Trajectory PredictionHao Zhou, Lu Qi, Xiangtai Li, Jie Zhang 等CVPR 2026 · 被引用 2 次
- UniMotion: A Unified Motion Framework for Simulation, Prediction and PlanningNan Song, Junzhe Jiang, Jingyu Li, Xiatian Zhu 等NeurIPS 2025 · 被引用 2 次
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- DenseTNT: End-to-end Trajectory Prediction from Dense Goal SetsJunru Gu, Chen Sun, Hang ZhaoICCV 2021 · 被引用 563 次
- Motion Transformer with Global Intention Localization and Local Movement RefinementShaoshuai Shi, Li Jiang, Dengxin Dai, Bernt SchieleNeurIPS 2022 · 被引用 515 次
- HiVT: Hierarchical Vector Transformer for Multi-Agent Motion PredictionZikang Zhou, Luyao Ye, Jianping Wang, Kui Wu 等CVPR 2022 · 被引用 379 次
- Scene Transformer: A unified architecture for predicting future trajectories of multiple agentsJiquan Ngiam, Vijay Vasudevan, Benjamin Caine, Zhengdong Zhang 等ICLR 2022 · 被引用 194 次
相关 Paper
- Forecast-MAE: Self-supervised Pre-training for Motion Forecasting with Masked AutoencodersJie Cheng, Xiaodong Mei, Ming LiuICCV 2023 · 被引用 123 次
- Query-Centric Trajectory PredictionZikang Zhou, Jianping Wang, Yung-Hui Li, Yu-Kai HuangCVPR 2023
- DONUT: A Decoder-Only Model for Trajectory PredictionMarkus Knoche, Daan de Geus, Bastian LeibeICCV 2025
- LaPred: Lane-Aware Prediction of Multi-Modal Future Trajectories of Dynamic AgentsByeoungdo Kim, SeongHyeon Park, Seokhwan Lee, Elbek Khoshimjonov 等CVPR 2021
- Future-Aware Interaction Network for Motion ForecastingShijie Li, Chunyu Liu, Xun Xu, Si Yong Yeo 等ICCV 2025 · 被引用 4 次
