Social-Transmotion: Promptable Human Trajectory Prediction
Saeed Saadatnejad, Yang Gao, Kaouther Messaoud, Alexandre Alahi
Abstract
Accurate human trajectory prediction is crucial for applications such as autonomous vehicles, robotics, and surveillance systems. Yet, existing models often fail to fully leverage the non-verbal social cues human subconsciously communicate when navigating the space. To address this, we introduce Social-Transmotion, a generic Transformer-based model that exploits diverse and numerous visual cues to predict human behavior. We translate the idea of a prompt from Natural Language Processing (NLP) to the task of human trajectory prediction, where a prompt can be a sequence of x-y coordinates on the ground, bounding boxes in the image plane, or body pose keypoints in either 2D or 3D. This, in turn, augments trajectory data, leading to enhanced human trajectory prediction. Using masking technique, our model exhibits flexibility and adaptability by capturing spatiotemporal interactions between agents based on the available visual cues. We delve into the merits of using 2D versus 3D poses, and a limited set of poses. Additionally, we investigate the spatial and temporal attention map to identify which keypoints and time-steps in the sequence are vital for optimizing human trajectory prediction. Our approach is validated on multiple datasets, including JTA, JRDB, Pedestrians and Cyclists in Road Traffic, and ETH-UCY. The code is publicly available: https://github.com/vita-epfl/social-transmotion .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8f03df42-2f83-4a0f-ac6a-09595af88cc0Cited by top-tier papers11
- JointDiff: Bridging Continuous and Discrete in Multi-Agent Trajectory GenerationGuillem Capellera, Luis Ferraz, Antonio Romano, Alexandre Alahi et al.ICLR 2026 · 6 citations
- Deformable Gaussian Occupancy: Decoupling Rigid and Nonrigid Motion with Factorized DistillationYang Gao, Wuyang Li, Po-Chien Luan, Alexandre AlahiCVPR 2026 · 2 citations
- HUMOF: Human Motion Forecasting in Interactive Social ScenesCaiyi Sun, Yujing Sun, Xiao Han, Zemin Yang et al.ICLR 2026 · 2 citations
- Towards Predicting Any Human Trajectory In ContextRyo Fujii, Hideo Saito, Ryo HachiumaNeurIPS 2025 · 2 citations
- Enhancing Trajectory Prediction through Self-Supervised Waypoint Distortion PredictionPranav Singh Chib, Pravendra SinghICML 2024 · 2 citations
Builds on14
- AgentFormer: Agent-Aware Transformers for Socio-Temporal Multi-Agent ForecastingYe Yuan, Xinshuo Weng, Yanglan Ou, Kris KitaniICCV 2021 · 658 citations
- STGAT: Modeling Spatial-Temporal Interactions for Human Trajectory PredictionYingfan Huang, Huikun Bi, Zhaoxin Li, Tianlu Mao et al.ICCV 2019 · 615 citations
- From Goals, Waypoints & Paths To Long Term Human Trajectory ForecastingKarttikeya Mangalam, Yang An, Harshayu Girase, Jitendra MalikICCV 2021 · 345 citations
- Stochastic Trajectory Prediction via Motion Indeterminacy DiffusionTianpei Gu, Guangyi Chen, Junlong Li, Chunze Lin et al.CVPR 2022 · 261 citations
- Latent Variable Sequential Set Transformers for Joint Multi-Agent Motion PredictionRoger Girgis, Florian Golemo, Felipe Codevilla, Martin Weiss et al.ICLR 2022 · 200 citations
Related papers
- Multi-Person 3D Motion Prediction with Multi-Range TransformersJiashun Wang, Huazhe Xu, Medhini Narasimhan, Xiaolong WangNeurIPS 2021 · 102 citations
- Introvert: Human Trajectory Prediction via Conditional 3D AttentionNasim Shafiee, Taskin Padir, Ehsan ElhamifarCVPR 2021
- Multimodal Interaction-Aware Trajectory Prediction in Crowded SpaceXiaodan Shi, Xiaowei Shao, Zipei Fan, Renhe Jiang et al.AAAI 2020 · 32 citations
- RealisMotion: Decomposed Human Motion Control and Video Generation in the World SpaceJingyun Liang, Jingkai Zhou, Shikai Li, Chenjie Cao et al.ICML 2026 · 9 citations
- Motion Transformer with Global Intention Localization and Local Movement RefinementShaoshuai Shi, Li Jiang, Dengxin Dai, Bernt SchieleNeurIPS 2022 · 515 citations
