Unified Multi-Agent Trajectory Modeling with Masked Trajectory Diffusion
Songru Yang, Zhenwei Shi, Zhengxia Zou
Abstract
Understanding movements in multi-agent scenarios is a fundamental problem in intelligent systems. Previous research assumes complete and synchronized observations. However, real-world partial observation caused by occlusions leads to inevitable model failure, which demands a unified framework for coexisting trajectory prediction, imputation, and recovery. Unlike previous attempts that handled observed and unobserved behaviors in a coupled manner, we explore a decoupled denoising diffusion modeling paradigm with a unidirectional information valve to separate the interference from uncertain behaviors. Building on this, we proposed a Unified Masked Trajectory Diffusion model (UniMTD) for arbitrary levels of missing observations. We designed a unidirectional attention as a valve unit to control the direction of information flow between the observed and masked areas, gradually refining the missing observations toward a real-world distribution. We construct it into a unidirectional MoE structure to handle varying proportions of missing observations. A Cached Diffusion model is further designed to improve generation quality while reducing computation and time overhead. Our method has achieved a great leap across human motions and vehicle traffic. UniMTD efficiently achieves 74% improvement in minADE 20 and reaches SOTA with advantages of 91%, 66%, 69%, and 58% across 4 fidelity metrics on out-of-boundary, velocity, and trajectory length. (a) UniMTD UniTraj GC-VRNN SSSD INAM Naomi MAT Transformer LSTM (b)
… Mixed Encoder Masked Observed Trajectory with arbitrary observation loss Noised Latent Space Vanilla Decoder Low-quality results Previous Methods UniMTD(Ours) Accurate Latent Space Observed Behaviors Preservation Masked Behaviors Projection Efficient Cached Diffusion Model Unidirectional Encoder High-quality results Reuse Calculation (c) Real-world Distribution
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- Autoregressive Denoising Diffusion Models for Multivariate Probabilistic Time Series ForecastingKashif Rasul, Calvin Seward, Ingmar Schuster, Roland VollgrafICML 2021 · 500 citations
- HiVT: Hierarchical Vector Transformer for Multi-Agent Motion PredictionZikang Zhou, Luyao Ye, Jianping Wang, Kui Wu et al.CVPR 2022 · 379 citations
Related papers
- A Universal Model for Human Mobility PredictionQingyue Long, Yuan Yuan, Yong LiKDD 2025 · 8 citations
- Unified Uncertainty-Aware Diffusion for Multi-Agent Trajectory ModelingGuillem Capellera, Antonio Rubio, Luis Ferraz, Antonio AgudoCVPR 2025
- Sports-Traj: A Unified Trajectory Generation Model for Multi-Agent Movement in SportsYi Xu, Yun FuICLR 2025
- Uncovering the Missing Pattern: Unified Framework Towards Trajectory Imputation and PredictionYi Xu, Armin Bazarjani, Hyung-Gun Chi, Chiho Choi et al.CVPR 2023
- UniMotion: A Unified Motion Framework for Simulation, Prediction and PlanningNan Song, Junzhe Jiang, Jingyu Li, Xiatian Zhu et al.NeurIPS 2025 · 2 citations
