Policy Pre-training for Autonomous Driving via Self-supervised Geometric Modeling
Penghao Wu, Li Chen, Hongyang Li, Xiaosong Jia, Junchi Yan, Yu Qiao
Abstract
Witnessing the impressive achievements of pre-training techniques on large-scale data in the field of computer vision and natural language processing, we wonder whether this idea could be adapted in a grab-and-go spirit, and mitigate the sample inefficiency problem for visuomotor driving. Given the highly dynamic and variant nature of the input, the visuomotor driving task inherently lacks view and translation invariance, and the visual input contains massive irrelevant information for decision making, resulting in predominant pre-training approaches from general vision less suitable for the autonomous driving task. To this end, we propose PPGeo (Policy Pre-training via Geometric modeling), an intuitive and straightforward fully self-supervised framework curated for the policy pretraining in visuomotor driving. We aim at learning policy representations as a powerful abstraction by modeling 3D geometric scenes on large-scale unlabeled and uncalibrated YouTube driving videos. The proposed PPGeo is performed in two stages to support effective self-supervised training. In the first stage, the geometric modeling framework generates pose and depth predictions simultaneously, with two consecutive frames as input. In the second stage, the visual encoder learns driving policy representation by predicting the future ego-motion and optimizing with the photometric error based on current visual observation only. As such, the pre-trained visual encoder is equipped with rich driving policy related representations and thereby competent for multiple visuomotor driving tasks. As a side product, the pre-trained geometric modeling networks could bring further improvement to the depth and odometry estimation tasks. Extensive experiments covering a wide span of challenging scenarios have demonstrated the superiority of our proposed approach, where improvements range from 2% to even over 100% with very limited data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bb6d9dc9-ca93-4e63-bb92-c13968020563Cited by top-tier papers11
- Vista: A Generalizable Driving World Model with High Fidelity and Versatile ControllabilityShenyuan Gao, Jiazhi Yang, Li Chen, Kashyap Chitta et al.NeurIPS 2024 · 403 citations
- DriveAdapter: Breaking the Coupling Barrier of Perception and Planning in End-to-End Autonomous DrivingXiaosong Jia, Yulu Gao, Li Chen, Junchi Yan et al.ICCV 2023 · 154 citations
- Raw2Drive: Reinforcement Learning with Aligned World Models for End-to-End Autonomous Driving (in CARLA v2)Zhenjie Yang, Xiaosong Jia, Qifeng Li, Xue Yang et al.NeurIPS 2025 · 65 citations
- ReSim: Reliable World Simulation for Autonomous DrivingJiazhi Yang, Kashyap Chitta, Shenyuan Gao, Long Chen et al.NeurIPS 2025 · 53 citations
- Visual Point Cloud Forecasting Enables Scalable Autonomous DrivingZetong Yang, Li Chen, Yanan Sun, Hongyang LiCVPR 2024 · 40 citations
Builds on25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- SimMIM: a Simple Framework for Masked Image ModelingZhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin et al.CVPR 2022 · 1,129 citations
Related papers
- Learning to Drive is a Free Gift: Large-Scale Label-Free Autonomy Pretraining from Unposed In-The-Wild VideosMatthew Strong, Wei-Jer Chang, Quentin Herau, Jiezhi Yang et al.CVPR 2026 · 3 citations
- Reinforcement Learning with Action-Free Pre-Training from VideosYounggyo Seo, Kimin Lee, Stephen James, Pieter AbbeelICML 2022 · 150 citations
- LA-Pose: Latent Action Pretraining Meets Pose EstimationZhengqing Wang, Saurabh Nair, Prajwal Chidananda, Pujith Kachana et al.CVPR 2026
- Self-Supervised Pretraining for Large-Scale Point CloudsZaiwei Zhang, Min Bai, Li Erran LiNeurIPS 2022 · 12 citations
- SelfD: Self-Learning Large-Scale Driving Policies From the WebJimuyang Zhang, Ruizhao Zhu, Eshed Ohn-BarCVPR 2022 · 17 citations
