Closed-Loop Supervised Fine-Tuning of Tokenized Traffic Models
Zhejun Zhang, Péter Karkus, Maximilian Igl, Wenhao Ding, Yuxiao Chen, Boris Ivanovic, Marco Pavone
Abstract
Traffic simulation aims to learn a policy for traffic agents that, when unrolled in closed-loop, faithfully recovers the joint distribution of trajectories observed in the real world. Inspired by large language models, tokenized multi-agent policies have recently become the state-of-the-art in traffic simulation. However, they are typically trained through open-loop behavior cloning, and thus suffer from covariate shift when executed in closed-loop during simulation. In this work, we present Closest Among Top-K (CAT-K) rollouts, a simple yet effective closed-loop fine-tuning strategy to mitigate covariate shift. CAT-K fine-tuning only requires existing trajectory data, without reinforcement learning or generative adversarial imitation. Concretely, CAT-K finetuning enables a small 7M-parameter tokenized traffic simulation policy to outperform a 102M-parameter model from the same model family, achieving the top spot on the Waymo Sim Agent Challenge leaderboard at the time of submission. The code is available at https://github.com/ NVlabs/catk.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 981a6457-e978-4d3b-b10d-692303b8f9d7Cited by top-tier papers12
- AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-TuningZewei Zhou, Tianhui Cai, Seth Z. Zhao, Yun Zhang et al.NeurIPS 2025 · 310 citations
- Advancing Multi-agent Traffic Simulation via R1-Style Reinforcement Fine-TuningMuleilan Pei, Shaoshuai Shi, Shaojie ShenICLR 2026 · 21 citations
- SPACeR: Self-Play Anchoring with Centralized Reference ModelsWei-Jer Chang, Akshay Rangesh, Kevin Joseph, Matthew Strong et al.ICLR 2026 · 9 citations
- DecompGAIL: Learning Realistic Traffic Behaviors with Decomposed Multi-Agent Generative Adversarial Imitation LearningKe Guo, Haochen Liu, Xiaojun Wu, Chen LvICLR 2026 · 8 citations
- LangTraj: Diffusion Model and Dataset for Language-Conditioned Trajectory SimulationWei-Jer Chang, Wei Zhan, Masayoshi Tomizuka, Manmohan Chandraker et al.ICCV 2025 · 5 citations
Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Large Scale Interactive Motion Forecasting for Autonomous Driving : The Waymo Open Motion DatasetScott Ettinger, Shuyang Cheng, Benjamin Caine, Chenxi Liu et al.ICCV 2021 · 817 citations
- Motion Transformer with Global Intention Localization and Local Movement RefinementShaoshuai Shi, Li Jiang, Dengxin Dai, Bernt SchieleNeurIPS 2022 · 515 citations
- The Trajectron: Probabilistic Multi-Agent Trajectory Modeling With Dynamic Spatiotemporal GraphsBoris Ivanovic, Marco PavoneICCV 2019 · 473 citations
Related papers
- SceneDiffuser: Efficient and Controllable Driving Simulation Initialization and RolloutChiyu Max Jiang, Yijing Bai, Andre Cornman, Christopher Davis et al.NeurIPS 2024 · 76 citations
- RLFTSim: Realistic and Controllable Multi-Agent Traffic Simulation via Reinforcement Learning Fine-TuningEhsan Ahmadi, Hunter Schofield, Behzad Khamidehi, Fazel Arasteh et al.CVPR 2026 · 2 citations
- BehaviorGPT: Smart Agent Simulation for Autonomous Driving with Next-Patch PredictionZikang Zhou, Haibo Hu, Xinhong Chen, Jianping Wang et al.NeurIPS 2024 · 73 citations
- MotionLM: Multi-Agent Motion Forecasting as Language ModelingAri Seff, Brian Cera, Dian Chen, Mason Ng et al.ICCV 2023 · 186 citations
- Prompt to Transfer: Sim-to-Real Transfer for Traffic Signal Control with Prompt LearningLongchao Da, Minquan Gao, Hao Mei, Hua WeiAAAI 2024 · 60 citations
