Robust Autonomy Emerges from Self-Play
Marco Francis Cusumano-Towner, David Hafner, Alexander Hertzberg, Brody Huval, Aleksei Petrenko, Eugene Vinitsky, Erik Wijmans, Taylor W. Killian, Stuart Bowers, Ozan Sener, Philipp Krähenbühl, Vladlen Koltun
Abstract
Self-play has powered breakthroughs in twoplayer and multi-player games. Here we show that self-play is a surprisingly effective strategy in another domain. We show that robust and naturalistic driving emerges entirely from self-play in simulation at unprecedented scale -1.6 billion km of driving. This is enabled by GIGAFLOW, a batched simulator that can synthesize and train on 42 years of subjective driving experience per hour on a single 8-GPU node. The resulting policy achieves state-of-the-art performance on three independent autonomous driving benchmarks. The policy outperforms the prior state of the art when tested on recorded real-world scenarios, amidst human drivers, without ever seeing human data during training. The policy is realistic when assessed against human references and achieves unprecedented robustness, averaging 17.5 years of continuous driving between incidents in simulation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b2d3c9ee-22b9-4767-aab0-1178afa56b37Cited by top-tier papers13
- ReSim: Reliable World Simulation for Autonomous DrivingJiazhi Yang, Kashyap Chitta, Shenyuan Gao, Long Chen et al.NeurIPS 2025 · 53 citations
- SimScale: Learning to Drive via Real-World Simulation at ScaleHaochen Tian, Tianyu Li, Haochen Liu, Jiazhi Yang et al.CVPR 2026 · 40 citations
- Flow Matching-Based Autonomous Driving Planning with Advanced Interactive Behavior ModelingTianyi Tan, Yinan Zheng, Ruiming Liang, Zexu Wang et al.NeurIPS 2025 · 36 citations
- LEAD: Minimizing Learner-Expert Asymmetry in End-to-End DrivingLong Nguyen, Micha Fauth, Bernhard Jaeger, Daniel Dauner et al.CVPR 2026 · 28 citations
- Plan-R1: Safe and Feasible Trajectory Planning as Language ModelingXiaolong Tang, Meina Kan, Shiguang Shan, Xilin ChenICLR 2026 · 26 citations
Builds on13
- Large Scale Interactive Motion Forecasting for Autonomous Driving : The Waymo Open Motion DatasetScott Ettinger, Shuyang Cheng, Benjamin Caine, Chenxi Liu et al.ICCV 2021 · 817 citations
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein et al.ASPLOS 2024 · 693 citations
- DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion FramesErik Wijmans, Abhishek Kadian, Ari Morcos, Stefan Lee et al.ICLR 2020 · 608 citations
- Motion Transformer with Global Intention Localization and Local Movement RefinementShaoshuai Shi, Li Jiang, Dengxin Dai, Bernt SchieleNeurIPS 2022 · 515 citations
- The Trajectron: Probabilistic Multi-Agent Trajectory Modeling With Dynamic Spatiotemporal GraphsBoris Ivanovic, Marco PavoneICCV 2019 · 473 citations
Related papers
- SPACeR: Self-Play Anchoring with Centralized Reference ModelsWei-Jer Chang, Akshay Rangesh, Kevin Joseph, Matthew Strong et al.ICLR 2026 · 9 citations
- Uncertainty-Guided Never-Ending Learning to DriveLei Lai, Eshed Ohn-Bar, Sanjay Arora, John Seon Keun YiCVPR 2024
- Deformation and Correspondence Aware Unsupervised Synthetic-to-Real Scene Flow Estimation for Point CloudsZhao Jin, Yinjie Lei, Naveed Akhtar, Haifeng Li et al.CVPR 2022 · 26 citations
- Model-Based Imitation Learning for Urban DrivingAnthony Hu, Gianluca Corrado, Nicolas Griffiths, Zachary Murez et al.NeurIPS 2022 · 241 citations
- Learning from All VehiclesDian Chen, Philipp KrähenbühlCVPR 2022
