Learning to Drive is a Free Gift: Large-Scale Label-Free Autonomy Pretraining from Unposed In-The-Wild Videos
Matthew Strong, Wei-Jer Chang, Quentin Herau, Jiezhi Yang, Yihan Hu, Chensheng Peng, Wei Zhan
摘要
Ego-centric driving videos available online provide an abundant source of visual data for autonomous driving, yet their lack of annotations makes it difficult to learn representations that capture both semantic structure and 3D geometry. Recent advances in large feedforward spatial models demonstrate that point maps and ego-motion can be inferred in a single forward pass, suggesting a promising direction for scalable driving perception. We therefore propose a label-free, teacher-guided framework for learning autonomous driving representations directly from unposed videos. Unlike prior self-supervised approaches that focus primarily on frame-to-frame consistency, we posit that safe and reactive driving depends critically on temporal context. To this end, we leverage a feedforward architecture equipped with a lightweight autoregressive module, trained using multi-modal supervisory signals that guide the model to jointly predict current and future point maps, camera poses, semantic segmentation, and motion masks. Multi-modal teachers provide sequence-level pseudo-supervision, enabling LFG to learn a unified pseudo-4D representation from raw YouTube videos without poses, labels, or LiDAR. The resulting encoder not only transfers effectively to downstream autonomous driving planning on the NAVSIM benchmark, surpassing multi-camera and LiDAR baselines with only a single monocular camera, but also yields strong performance when evaluated on a range of semantic, geometric, and qualitative motion prediction tasks. These geometry and motion-aware features position LFG as a compelling video-centric foundation model for autonomous driving.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
- π3: Permutation-Equivariant Visual Geometry LearningYifan Wang, Jianjun Zhou, Haoyi Zhu, Wenzheng Chang 等ICLR 2026 · 被引用 318 次
- DrivingForward: Feed-forward 3D Gaussian Splatting for Driving Scene Reconstruction from Flexible Surround-view InputQijian Tian, Xin Tan, Yuan Xie, Lizhuang MaAAAI 2025 · 被引用 45 次
- Visual Point Cloud Forecasting Enables Scalable Autonomous DrivingZetong Yang, Li Chen, Yanan Sun, Hongyang LiCVPR 2024 · 被引用 40 次
相关 Paper
- Policy Pre-training for Autonomous Driving via Self-supervised Geometric ModelingPenghao Wu, Li Chen, Hongyang Li, Xiaosong Jia 等ICLR 2023 · 被引用 7 次
- Enhancing End-to-End Autonomous Driving with Latent World ModelYingyan Li, Lue Fan, Jiawei He, Yuqi Wang 等ICLR 2025
- World4Drive: End-to-End Autonomous Driving via Intention-Aware Physical Latent World ModelYupeng Zheng, Pengxuan Yang, Zebin Xing, Qichao Zhang 等ICCV 2025 · 被引用 16 次
- TrajVG: 3D Trajectory-Coupled Visual Geometry LearningXingyu Miao, Weiguang Zhao, Tao Lu, Linning Xu 等SIGGRAPH 2026
- SkyEye: Self-Supervised Bird's-Eye-View Semantic Mapping Using Monocular Frontal View ImagesNikhil Gosala, Kürsat Petek, Paulo L. J. Drews-Jr, Wolfram Burgard 等CVPR 2023
