MotionLM: Multi-Agent Motion Forecasting as Language Modeling
Ari Seff, Brian Cera, Dian Chen, Mason Ng, Aurick Zhou, Nigamaa Nayakanti, Khaled S. Refaat, Rami Al-Rfou, Benjamin Sapp
Abstract
Reliable forecasting of the future behavior of road agents is a critical component to safe planning in autonomous vehicles. Here, we represent continuous trajectories as sequences of discrete motion tokens and cast multi-agent motion prediction as a language modeling task over this domain. Our model, MotionLM, provides several advantages: First, it does not require anchors or explicit latent variable optimization to learn multimodal distributions. Instead, we leverage a single standard language modeling objective, maximizing the average log probability over sequence tokens. Second, our approach bypasses post-hoc interaction heuristics where individual agent trajectory generation is conducted prior to interactive scoring. Instead, MotionLM produces joint distributions over interactive agent futures in a single autoregressive decoding process. In addition, the model's sequential factorization enables temporally causal conditional rollouts. The proposed approach establishes new state-of-the-art performance for multi-agent motion prediction on the Waymo Open Motion Dataset, ranking 1 st on the interactive challenge leaderboard.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ea79aade-c2c6-4177-a763-a7aa47fedba2Cited by top-tier papers41
- VADv2: End-to-End Vectorized Autonomous Driving via Probabilistic PlanningBo Jiang, Shaoyu Chen, Hao Gao, Bencheng Liao et al.ICLR 2026 · 259 citations
- SMART: Scalable Multi-agent Real-time Motion Generation via Next-token PredictionWei Wu, Xiaoxin Feng, Ziyan Gao, Yuheng KanNeurIPS 2024 · 104 citations
- SceneDiffuser: Efficient and Controllable Driving Simulation Initialization and RolloutChiyu Max Jiang, Yijing Bai, Andre Cornman, Christopher Davis et al.NeurIPS 2024 · 76 citations
- BehaviorGPT: Smart Agent Simulation for Autonomous Driving with Next-Patch PredictionZikang Zhou, Haibo Hu, Xinhong Chen, Jianping Wang et al.NeurIPS 2024 · 73 citations
- HPNet: Dynamic Trajectory Forecasting with Historical Prediction AttentionXiaolong Tang, Meina Kan, Shiguang Shan, Zhilong Ji et al.CVPR 2024 · 66 citations
Builds on11
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Large Scale Interactive Motion Forecasting for Autonomous Driving : The Waymo Open Motion DatasetScott Ettinger, Shuyang Cheng, Benjamin Caine, Chenxi Liu et al.ICCV 2021 · 817 citations
- AgentFormer: Agent-Aware Transformers for Socio-Temporal Multi-Agent ForecastingYe Yuan, Xinshuo Weng, Yanglan Ou, Kris KitaniICCV 2021 · 658 citations
- Motion Transformer with Global Intention Localization and Local Movement RefinementShaoshuai Shi, Li Jiang, Dengxin Dai, Bernt SchieleNeurIPS 2022 · 515 citations
Related papers
- Scene Transformer: A unified architecture for predicting future trajectories of multiple agentsJiquan Ngiam, Vijay Vasudevan, Benjamin Caine, Zhengdong Zhang et al.ICLR 2022 · 194 citations
- Trajeglish: Traffic Modeling as Next-Token PredictionJonah Philion, Xue Bin Peng, Sanja FidlerICLR 2024 · 61 citations
- M2I: From Factored Marginal Trajectory Prediction to Interactive PredictionQiao Sun, Xin Huang, Junru Gu, Brian C. Williams et al.CVPR 2022 · 104 citations
- DrivingGPT: Unifying Driving World Modeling and Planning with Multi-Modal Autoregressive TransformersYuntao Chen, Yuqi Wang, Zhaoxiang ZhangICCV 2025 · 7 citations
- DONUT: A Decoder-Only Model for Trajectory PredictionMarkus Knoche, Daan de Geus, Bastian LeibeICCV 2025
