AgentFormer: Agent-Aware Transformers for Socio-Temporal Multi-Agent Forecasting
Ye Yuan, Xinshuo Weng, Yanglan Ou, Kris Kitani
Abstract
Predicting accurate future trajectories of multiple agents is essential for autonomous systems but is challenging due to the complex interaction between agents and the uncertainty in each agent’s future behavior. Forecasting multi-agent trajectories requires modeling two key dimensions: (1) time dimension, where we model the influence of past agent states over future states; (2) social dimension, where we model how the state of each agent affects others. Most prior methods model these two dimensions separately, e.g., first using a temporal model to summarize features over time for each agent independently and then modeling the interaction of the summarized features with a social model. This approach is suboptimal since independent feature encoding over either the time or social dimension can result in a loss of information. Instead, we would prefer a method that allows an agent’s state at one time to directly affect another agent’s state at a future time. To this end, we propose a new Transformer, termed AgentFormer, that simultaneously models the time and social dimensions. The model leverages a sequence representation of multi-agent trajectories by flattening trajectory features across time and agents. Since standard attention operations disregard the agent identity of each element in the sequence, AgentFormer uses a novel agent-aware attention mechanism that preserves agent identities by attending to elements of the same agent differently than elements of other agents. Based on AgentFormer, we propose a stochastic multi-agent trajectory prediction model that can attend to features of any agent at any previous timestep when inferring an agent’s future position. The latent intent of all agents is also jointly modeled, allowing the stochasticity in one agent’s behavior to affect other agents. Extensive experiments show that our method significantly improves the state of the art on well-established pedestrian and autonomous driving datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 43bb385b-8770-4717-a814-613fb7dec9ceCited by top-tier papers91
- HiVT: Hierarchical Vector Transformer for Multi-Agent Motion PredictionZikang Zhou, Luyao Ye, Jianping Wang, Kui Wu et al.CVPR 2022 · 379 citations
- GameFormer: Game-theoretic Modeling and Learning of Transformer-based Interactive Prediction and Planning for Autonomous DrivingZhiyu Huang, Haochen Liu, Chen LvICCV 2023 · 209 citations
- Latent Variable Sequential Set Transformers for Joint Multi-Agent Motion PredictionRoger Girgis, Florian Golemo, Felipe Codevilla, Martin Weiss et al.ICLR 2022 · 200 citations
- MotionLM: Multi-Agent Motion Forecasting as Language ModelingAri Seff, Brian Cera, Dian Chen, Mason Ng et al.ICCV 2023 · 186 citations
- THOMAS: Trajectory Heatmap Output with learned Multi-Agent SamplingThomas Gilles, Stefano Sabatini, Dzmitry Tsishkou, Bogdan Stanciulescu et al.ICLR 2022 · 184 citations
Builds on14
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- STGAT: Modeling Spatial-Temporal Interactions for Human Trajectory PredictionYingfan Huang, Huikun Bi, Zhaoxin Li, Tianlu Mao et al.ICCV 2019 · 615 citations
- The Trajectron: Probabilistic Multi-Agent Trajectory Modeling With Dynamic Spatiotemporal GraphsBoris Ivanovic, Marco PavoneICCV 2019 · 473 citations
- PRECOG: PREdiction Conditioned on Goals in Visual Multi-Agent SettingsNicholas Rhinehart, Rowan McAllister, Kris Kitani, Sergey LevineICCV 2019 · 407 citations
Related papers
- Multimodal Interaction-Aware Trajectory Prediction in Crowded SpaceXiaodan Shi, Xiaowei Shao, Zipei Fan, Renhe Jiang et al.AAAI 2020 · 32 citations
- Multi-Person 3D Motion Prediction with Multi-Range TransformersJiashun Wang, Huazhe Xu, Medhini Narasimhan, Xiaolong WangNeurIPS 2021 · 102 citations
- Scene Transformer: A unified architecture for predicting future trajectories of multiple agentsJiquan Ngiam, Vijay Vasudevan, Benjamin Caine, Zhengdong Zhang et al.ICLR 2022 · 194 citations
- Trajectory Unified Transformer for Pedestrian Trajectory PredictionLiushuai Shi, Le Wang, Sanping Zhou, Gang HuaICCV 2023 · 100 citations
- TGFormer: Transformer with Track Query Group for Multi-Object TrackingRui Zeng, Yuanzhou Huang, Songwei PeiAAAI 2025 · 6 citations
