SelfD: Self-Learning Large-Scale Driving Policies From the Web
Jimuyang Zhang, Ruizhao Zhu, Eshed Ohn-Bar
Abstract
Effectively utilizing the vast amounts of ego-centric navigation data that is freely available on the internet can advance generalized intelligent systems, i.e., to robustly scale across perspectives, platforms, environmental conditions, scenarios, and geographical locations. However, it is difficult to directly leverage such large amounts of unlabeled and highly diverse datafor complex 3D reasoning and planning tasks. Consequently, researchers have primarily focused on its use for various auxiliary pixel- and image-level computer vision tasks that do not consider an ultimate navigational objective. In this work, we introduce SelfD, a framework for learning scalable driving by utilizing large amounts of online monocular images. Our key idea is to leverage iterative semi-supervised training when learning imitative agents from unlabeled data. To handle unconstrained viewpoints, scenes, and camera parameters, we train an image-based model that directly learns to plan in the Bird's Eye View (BEV) space. Next, we use unla-beled data to augment the decision-making knowledge and robustness of an initially trained model via self-training. In particular, we propose a pseudo-labeling step which enables making full use of highly diverse demonstration data through “hypothetical” planning-based data augmentation. We employ a large dataset of publicly available YouTube videos to train SelfD and comprehensively analyze its generalization benefits across challenging navigation scenarios. Without requiring any additional data collection or annotation efforts, SelfD demonstrates consistent improvements (by up to 24%) in driving performance evaluation on nuScenes, Argoverse, Waymo, and CARLA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- XVO: Generalized Visual Odometry via Cross-Modal Self-TrainingLei Lai, Zhongkai Shangguan, Jimuyang Zhang, Eshed Ohn-BarICCV 2023 · 27 citations
- AdaptAgent: Adapting Multimodal Web Agents with Few-Shot Learning from Human DemonstrationsGaurav Verma, Rachneet Kaur, Nishan Srishankar, Zhen Zeng et al.ACL 2025 · 19 citations
- Feedback-Guided Autonomous DrivingJimuyang Zhang, Zanming Huang, Arijit Ray, Eshed Ohn-BarCVPR 2024 · 15 citations
- Policy Pre-training for Autonomous Driving via Self-supervised Geometric ModelingPenghao Wu, Li Chen, Hongyang Li, Xiaosong Jia et al.ICLR 2023 · 7 citations
- Learning to Drive is a Free Gift: Large-Scale Label-Free Autonomy Pretraining from Unposed In-The-Wild VideosMatthew Strong, Wei-Jer Chang, Quentin Herau, Jiezhi Yang et al.CVPR 2026 · 3 citations
Builds on26
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
- Exploring the Limitations of Behavior Cloning for Autonomous DrivingFelipe Codevilla, Eder Santana, Antonio M. López, Adrien GaidonICCV 2019 · 666 citations
- In Defense of Pseudo-Labeling: An Uncertainty-Aware Pseudo-label Selection Framework for Semi-Supervised LearningMamshad Nayeem Rizve, Kevin Duarte, Yogesh S. Rawat, Mubarak ShahICLR 2021 · 630 citations
- Scaling and Benchmarking Self-Supervised Visual Representation LearningPriya Goyal, Dhruv Mahajan, Abhinav Gupta, Ishan MisraICCV 2019 · 429 citations
Related papers
- SimScale: Learning to Drive via Real-World Simulation at ScaleHaochen Tian, Tianyu Li, Haochen Liu, Jiazhi Yang et al.CVPR 2026 · 40 citations
- Uncertainty-Guided Never-Ending Learning to DriveLei Lai, Eshed Ohn-Bar, Sanjay Arora, John Seon Keun YiCVPR 2024
- SkyEye: Self-Supervised Bird's-Eye-View Semantic Mapping Using Monocular Frontal View ImagesNikhil Gosala, Kürsat Petek, Paulo L. J. Drews-Jr, Wolfram Burgard et al.CVPR 2023
- Leveraging Temporal Cues for Semi-Supervised Multi-View 3D Object DetectionJinhyung Park, Navyata Sanghvi, Hiroki Adachi, Yoshihisa Shibata et al.CVPR 2025
- Learning from All VehiclesDian Chen, Philipp KrähenbühlCVPR 2022
