ModeSeq: Taming Sparse Multimodal Motion Prediction with Sequential Mode Modeling
Zikang Zhou, Hengjian Zhou, Haibo Hu, Zihao Wen, Jianping Wang, Yung-Hui Li, Yu-Kai Huang
摘要
Anticipating the multimodality of future events lays the foundation for safe autonomous driving. However, multimodal motion prediction for traffic agents has been clouded by the lack of multimodal ground truth. Existing works predominantly adopt the winner-take-all training strategy to tackle this challenge, yet still suffer from limited trajectory diversity and uncalibrated mode confidence. While some approaches address these limitations by generating excessive trajectory candidates, they necessitate a postprocessing stage to identify the most representative modes, a process lacking universal principles and compromising trajectory accuracy. We are thus motivated to introduce ModeSeq, a new multimodal prediction paradigm that models modes as sequences. Unlike the common practice of decoding multiple plausible trajectories in one shot, Mod-eSeq requires motion decoders to infer the next mode step by step, thereby more explicitly capturing the correlation between modes and significantly enhancing the ability to reason about multimodality. Leveraging the inductive bias of sequential mode prediction, we also propose the Early-Match-Take-All (EMTA) training strategy to diversify the trajectories further. Without relying on dense mode prediction or heuristic post-processing, ModeSeq considerably improves the diversity of multimodal output while attaining satisfactory trajectory accuracy, resulting in balanced performance on motion prediction benchmarks. Moreover, ModeSeq naturally emerges with the capability of mode extrapolation, which supports forecasting more behavior modes when the future is highly uncertain.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Large Scale Interactive Motion Forecasting for Autonomous Driving : The Waymo Open Motion DatasetScott Ettinger, Shuyang Cheng, Benjamin Caine, Chenxi Liu 等ICCV 2021 · 被引用 817 次
- Motion Transformer with Global Intention Localization and Local Movement RefinementShaoshuai Shi, Li Jiang, Dengxin Dai, Bernt SchieleNeurIPS 2022 · 被引用 515 次
- PRECOG: PREdiction Conditioned on Goals in Visual Multi-Agent SettingsNicholas Rhinehart, Rowan McAllister, Kris Kitani, Sergey LevineICCV 2019 · 被引用 407 次
- HiVT: Hierarchical Vector Transformer for Multi-Agent Motion PredictionZikang Zhou, Luyao Ye, Jianping Wang, Kui Wu 等CVPR 2022 · 被引用 379 次
- MotionLM: Multi-Agent Motion Forecasting as Language ModelingAri Seff, Brian Cera, Dian Chen, Mason Ng 等ICCV 2023 · 被引用 186 次
相关 Paper
- DeMo: Decoupling Motion Forecasting into Directional Intentions and Dynamic StatesBozhou Zhang, Nan Song, Li ZhangNeurIPS 2024 · 被引用 34 次
- Multimodal Motion Prediction With Stacked TransformersYicheng Liu, Jinghuai Zhang, Liangji Fang, Qinhong Jiang 等CVPR 2021
- ProphNet: Efficient Agent-Centric Motion Forecasting with Anchor-Informed ProposalsXishun Wang, Tong Su, Fang Da, Xiaodong YangCVPR 2023
- TPNet: Trajectory Proposal Network for Motion PredictionLiangji Fang, Qinhong Jiang, Jianping Shi, Bolei ZhouCVPR 2020
- CoverNet: Multimodal Behavior Prediction Using Trajectory SetsTung Phan-Minh, Elena Corina Grigore, Freddy A. Boulton, Oscar Beijbom 等CVPR 2020
