Multimodal Motion Prediction With Stacked Transformers
Yicheng Liu, Jinghuai Zhang, Liangji Fang, Qinhong Jiang, Bolei Zhou
Abstract
Predicting multiple plausible future trajectories of the nearby vehicles is crucial for the safety of autonomous driving. Recent motion prediction approaches attempt to achieve such multimodal motion prediction by implicitly regularizing the feature or explicitly generating multiple candidate proposals. However, it remains challenging since the latent features may concentrate on the most frequent mode of the data while the proposal-based methods depend largely on the prior knowledge to generate and select the proposals. In this work, we propose a novel transformer framework for multimodal motion prediction, termed as mmTransformer. A novel network architecture based on stacked transformers is designed to model the multimodality at feature level with a set of fixed independent proposals. A region-based training strategy is then developed to induce the multimodality of the generated proposals. Experiments on Argoverse dataset show that the proposed model achieves the state-of-the-art performance on motion prediction, substantially improving the diversity and the accuracy of the predicted trajectories. Demo video and code are available at https://decisionforce . github.io/mmTransformer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f1b8efad-1ec8-4ce1-a654-f37e136798ffCited by top-tier papers65
- VAD: Vectorized Scene Representation for Efficient Autonomous DrivingBo Jiang, Shaoyu Chen, Qing Xu, Bencheng Liao et al.ICCV 2023 · 602 citations
- Motion Transformer with Global Intention Localization and Local Movement RefinementShaoshuai Shi, Li Jiang, Dengxin Dai, Bernt SchieleNeurIPS 2022 · 515 citations
- HiVT: Hierarchical Vector Transformer for Multi-Agent Motion PredictionZikang Zhou, Luyao Ye, Jianping Wang, Kui Wu et al.CVPR 2022 · 379 citations
- VectorMapNet: End-to-end Vectorized HD Map LearningYicheng Liu, Tianyuan Yuan, Yue Wang, Yilun Wang et al.ICML 2023 · 332 citations
- VADv2: End-to-End Vectorized Autonomous Driving via Probabilistic PlanningBo Jiang, Shaoyu Chen, Hao Gao, Bencheng Liao et al.ICLR 2026 · 259 citations
Builds on5
- Neural Turtle Graphics for Modeling City Road LayoutsHang Chu, Daiqing Li, David Acuna, Amlan Kar et al.ICCV 2019 · 93 citations
- A Novel Learning Framework for Sampling-Based Motion Planning in Autonomous DrivingYifan Zhang, Jinghuai Zhang, Jindi Zhang, Jianping Wang et al.AAAI 2020 · 24 citations
- VectorNet: Encoding HD Maps and Agent Dynamics From Vectorized RepresentationJiyang Gao, Chen Sun, Hang Zhao, Yi Shen et al.CVPR 2020
- CoverNet: Multimodal Behavior Prediction Using Trajectory SetsTung Phan-Minh, Elena Corina Grigore, Freddy A. Boulton, Oscar Beijbom et al.CVPR 2020
- TPNet: Trajectory Proposal Network for Motion PredictionLiangji Fang, Qinhong Jiang, Jianping Shi, Bolei ZhouCVPR 2020
Related papers
- LTP: Lane-based Trajectory Prediction for Autonomous DrivingJingke Wang, Tengju Ye, Ziqing Gu, Junbo ChenCVPR 2022 · 75 citations
- Motion Diversification NetworksHee Jae Kim, Eshed Ohn-BarCVPR 2024
- Future-Aware Interaction Network for Motion ForecastingShijie Li, Chunyu Liu, Xun Xu, Si Yong Yeo et al.ICCV 2025 · 4 citations
- Motion Forecasting in Continuous DrivingNan Song, Bozhou Zhang, Xiatian Zhu, Li ZhangNeurIPS 2024 · 33 citations
- Bootstrap Motion Forecasting With Self-Consistent ConstraintsMaosheng Ye, Jiamiao Xu, Xunnong Xu, Tengfei Wang et al.ICCV 2023 · 27 citations
