Robust Autonomy Emerges from Self-Play
Marco Francis Cusumano-Towner, David Hafner, Alexander Hertzberg, Brody Huval, Aleksei Petrenko, Eugene Vinitsky, Erik Wijmans, Taylor W. Killian, Stuart Bowers, Ozan Sener, Philipp Krähenbühl, Vladlen Koltun
摘要
Self-play has powered breakthroughs in twoplayer and multi-player games. Here we show that self-play is a surprisingly effective strategy in another domain. We show that robust and naturalistic driving emerges entirely from self-play in simulation at unprecedented scale -1.6 billion km of driving. This is enabled by GIGAFLOW, a batched simulator that can synthesize and train on 42 years of subjective driving experience per hour on a single 8-GPU node. The resulting policy achieves state-of-the-art performance on three independent autonomous driving benchmarks. The policy outperforms the prior state of the art when tested on recorded real-world scenarios, amidst human drivers, without ever seeing human data during training. The policy is realistic when assessed against human references and achieves unprecedented robustness, averaging 17.5 years of continuous driving between incidents in simulation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- ReSim: Reliable World Simulation for Autonomous DrivingJiazhi Yang, Kashyap Chitta, Shenyuan Gao, Long Chen 等NeurIPS 2025 · 被引用 53 次
- SimScale: Learning to Drive via Real-World Simulation at ScaleHaochen Tian, Tianyu Li, Haochen Liu, Jiazhi Yang 等CVPR 2026 · 被引用 40 次
- Flow Matching-Based Autonomous Driving Planning with Advanced Interactive Behavior ModelingTianyi Tan, Yinan Zheng, Ruiming Liang, Zexu Wang 等NeurIPS 2025 · 被引用 36 次
- LEAD: Minimizing Learner-Expert Asymmetry in End-to-End DrivingLong Nguyen, Micha Fauth, Bernhard Jaeger, Daniel Dauner 等CVPR 2026 · 被引用 28 次
- Plan-R1: Safe and Feasible Trajectory Planning as Language ModelingXiaolong Tang, Meina Kan, Shiguang Shan, Xilin ChenICLR 2026 · 被引用 26 次
它引用的顶会 Paper13
- Large Scale Interactive Motion Forecasting for Autonomous Driving : The Waymo Open Motion DatasetScott Ettinger, Shuyang Cheng, Benjamin Caine, Chenxi Liu 等ICCV 2021 · 被引用 817 次
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein 等ASPLOS 2024 · 被引用 693 次
- DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion FramesErik Wijmans, Abhishek Kadian, Ari Morcos, Stefan Lee 等ICLR 2020 · 被引用 608 次
- Motion Transformer with Global Intention Localization and Local Movement RefinementShaoshuai Shi, Li Jiang, Dengxin Dai, Bernt SchieleNeurIPS 2022 · 被引用 515 次
- The Trajectron: Probabilistic Multi-Agent Trajectory Modeling With Dynamic Spatiotemporal GraphsBoris Ivanovic, Marco PavoneICCV 2019 · 被引用 473 次
相关 Paper
- SPACeR: Self-Play Anchoring with Centralized Reference ModelsWei-Jer Chang, Akshay Rangesh, Kevin Joseph, Matthew Strong 等ICLR 2026 · 被引用 9 次
- Uncertainty-Guided Never-Ending Learning to DriveLei Lai, Eshed Ohn-Bar, Sanjay Arora, John Seon Keun YiCVPR 2024
- Deformation and Correspondence Aware Unsupervised Synthetic-to-Real Scene Flow Estimation for Point CloudsZhao Jin, Yinjie Lei, Naveed Akhtar, Haifeng Li 等CVPR 2022 · 被引用 26 次
- Model-Based Imitation Learning for Urban DrivingAnthony Hu, Gianluca Corrado, Nicolas Griffiths, Zachary Murez 等NeurIPS 2022 · 被引用 241 次
- Learning from All VehiclesDian Chen, Philipp KrähenbühlCVPR 2022
