Can Language Beat Numerical Regression? Language-Based Multimodal Trajectory Prediction
Inhwan Bae, Junoh Lee, Hae-Gon Jeon
摘要
Language models have demonstrated impressive ability in context understanding and generative performance. Inspired by the recent success of language foundation models, in this paper, we propose LMTraj (Language-based Multimodal Trajectory predictor), which recasts the trajectory prediction task into a sort of question-answering problem. Departing from traditional numerical regression models, which treat the trajectory coordinate sequence as continuous signals, we consider them as discrete signals like text prompts. Specially, we first transform an input space for the trajectory coordinate into the natural language space. Here, the entire timeseries trajectories of pedestrians are converted into a text prompt, and scene images are described as text information through image captioning. The transformed numerical and image data are then wrapped into the question-answering template for use in a language model. Next, to guide the language model in understanding and reasoning high-level knowledge, such as scene context and social relationships between pedestrians, we introduce an auxiliary multi-task question and answering. We then train a numerical tokenizer with the prompt data. We encourage the tokenizer to separate the integer and decimal parts well, and leverage it to capture correlations between the consecutive numbers in the language model. Lastly, we train the language model using the numerical tokenizer and all of the question-answer prompts. Here, we propose a beam-search-based most-likely prediction and a temperature-based multimodal prediction to implement both deterministic and stochastic inferences. Applying our LMTraj, we show that the language-based model can be a powerful pedestrian trajectory predictor, and outperforms existing numerical-based predictor methods. Extensive experiments show that our LMTraj can successfully understand social relationships and accurately extrapolate the multimodal futures on the public pedestrian trajectory prediction benchmark. Code is publicly available at https: //github.com/inhwanbae/LMTrajectory .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- NATRA: Noise-Agnostic Framework for Trajectory Prediction with Noisy ObservationsRongqing Li, Changsheng Li, Ruilin Lv, Yuhang Li 等ICCV 2025 · 被引用 3 次
- Generative Active Learning for Long-Tail Trajectory Prediction via Controllable Diffusion ModelDaehee Park, Monu Surana, Pranav Desai, Ashish Mehta 等ICCV 2025 · 被引用 2 次
- W2W: Language-Model-Based Trajectory Prediction with Reinforcement LearningZirui Xu, Biao Yang, Rongrong Ni, Zhongkai Zhou 等CVPR 2026 · 被引用 2 次
- TrajEvo: Trajectory Prediction Heuristics Design via LLM-driven EvolutionZhikai Zhao, Chuanbo Hua, Federico Berto, Kanghoon Lee 等AAAI 2026 · 被引用 2 次
- Resonance: Learning to Predict Social-Aware Pedestrian Trajectories as Co-VibrationsConghao Wong, Ziqian Zou, Beihao XiaICCV 2025 · 被引用 1 次
它引用的顶会 Paper63
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- CogView: Mastering Text-to-Image Generation via TransformersMing Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng 等NeurIPS 2021 · 被引用 1,026 次
- AgentFormer: Agent-Aware Transformers for Socio-Temporal Multi-Agent ForecastingYe Yuan, Xinshuo Weng, Yanglan Ou, Kris KitaniICCV 2021 · 被引用 658 次
- STGAT: Modeling Spatial-Temporal Interactions for Human Trajectory PredictionYingfan Huang, Huikun Bi, Zhaoxin Li, Tianlu Mao 等ICCV 2019 · 被引用 615 次
相关 Paper
- Traj-MLLM: Can Multimodal Large Language Models Reform Trajectory Data Mining?Shuo Liu, Di Yao, Yan Lin, Gao Cong 等KDD 2026 · 被引用 2 次
- AutoTraces: Autoregressive Trajectory Forecasting via Multimodal Large Language ModelsTeng Wang, Yanting Lu, Ruize WangCVPR 2026
- MotionLM: Multi-Agent Motion Forecasting as Language ModelingAri Seff, Brian Cera, Dian Chen, Mason Ng 等ICCV 2023 · 被引用 186 次
- CoMaPOI: A Collaborative Multi-Agent Framework for Next POI Prediction Bridging the Gap Between Trajectory and LanguageLin Zhong, Lingzhi Wang, Xu Yang, Qing LiaoSIGIR 2025 · 被引用 6 次
- OmniPT: Unleashing the Potential of Large Vision Language Models for Pedestrian Tracking and UnderstandingTeng Fu, Mengyang Zhao, Ke Niu, Kaixin Peng 等AAAI 2026
