Advancing Multi-agent Traffic Simulation via R1-Style Reinforcement Fine-Tuning
Muleilan Pei, Shaoshuai Shi, Shaojie Shen
摘要
Scalable and realistic simulation of multi-agent traffic behavior is critical for advancing autonomous driving technologies. Although existing data-driven simulators have made significant strides in this domain, they predominantly rely on supervised learning to align simulated distributions with real-world driving scenarios. A persistent challenge, however, lies in the distributional shift that arises between training and testing, which often undermines model generalization in unseen environments. To address this limitation, we propose SMART-R1, a novel R1-style reinforcement fine-tuning paradigm tailored for next-token prediction models to better align agent behavior with human preferences and evaluation metrics. Our approach introduces a metric-oriented policy optimization algorithm to improve distribution alignment and an iterative "SFT-RFT-SFT" post-training strategy that alternates between Supervised Fine-Tuning (SFT) and Reinforcement Fine-Tuning (RFT) to maximize performance gains. Extensive experiments on the large-scale Waymo Open Motion Dataset (WOMD) validate the effectiveness of this simple yet powerful R1-style training framework in enhancing foundation models. The results on the Waymo Open Sim Agents Challenge (WOSAC) showcase that SMART-R1 achieves state-of-the-art performance with an overall realism meta score of 0.7858, ranking first on the leaderboard at the time of submission.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper19
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Large Scale Interactive Motion Forecasting for Autonomous Driving : The Waymo Open Motion DatasetScott Ettinger, Shuyang Cheng, Benjamin Caine, Chenxi Liu 等ICCV 2021 · 被引用 817 次
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language ModelsWenxuan Huang, Bohan Jia, Shaosheng Cao, Zheyu Ye 等ICLR 2026 · 被引用 670 次
- MotionLM: Multi-Agent Motion Forecasting as Language ModelingAri Seff, Brian Cera, Dian Chen, Mason Ng 等ICCV 2023 · 被引用 186 次
相关 Paper
- RLFTSim: Realistic and Controllable Multi-Agent Traffic Simulation via Reinforcement Learning Fine-TuningEhsan Ahmadi, Hunter Schofield, Behzad Khamidehi, Fazel Arasteh 等CVPR 2026 · 被引用 2 次
- SMART: Scalable Multi-agent Real-time Motion Generation via Next-token PredictionWei Wu, Xiaoxin Feng, Ziyan Gao, Yuheng KanNeurIPS 2024 · 被引用 104 次
- Closed-Loop Supervised Fine-Tuning of Tokenized Traffic ModelsZhejun Zhang, Péter Karkus, Maximilian Igl, Wenhao Ding 等CVPR 2025
- SPACeR: Self-Play Anchoring with Centralized Reference ModelsWei-Jer Chang, Akshay Rangesh, Kevin Joseph, Matthew Strong 等ICLR 2026 · 被引用 9 次
- BehaviorGPT: Smart Agent Simulation for Autonomous Driving with Next-Patch PredictionZikang Zhou, Haibo Hu, Xinhong Chen, Jianping Wang 等NeurIPS 2024 · 被引用 73 次
