Self-Supervised Simultaneous Multi-Step Prediction of Road Dynamics and Cost Map
Elmira Amirloo Abolfathi, Mohsen Rohani, Ershad Banijamali, Jun Luo, Pascal Poupart
Abstract
While supervised learning is widely used for perception modules in conventional autonomous driving solutions, scalability is hindered by the huge amount of data labeling needed. In contrast, while end-to-end architectures do not require labeled data and are potentially more scalable, interpretability is sacrificed. We introduce a novel architecture that is trained in a fully self-supervised fashion for simultaneous multi-step prediction of space-time cost map and road dynamics. Our solution replaces the manually designed cost function for motion planning with a learned high dimensional cost map that is naturally interpretable and allows diverse contextual information to be integrated without manual data labeling. Experiments on real world driving data show that our solution leads to lower number of collisions and road violations in long planning horizons in comparison to baselines, demonstrating the feasibility of fully self-supervised prediction without sacrificing scalability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e03f5c4a-0710-4fbb-95b0-4c877734682bBuilds on2
- PRECOG: PREdiction Conditioned on Goals in Visual Multi-Agent SettingsNicholas Rhinehart, Rowan McAllister, Kris Kitani, Sergey LevineICCV 2019 · 407 citations
- Deep Imitative Models for Flexible Inference, Planning, and ControlNicholas Rhinehart, Rowan McAllister, Sergey LevineICLR 2020 · 159 citations
Related papers
- Enhancing End-to-End Autonomous Driving with Latent World ModelYingyan Li, Lue Fan, Jiawei He, Yuqi Wang et al.ICLR 2025
- World4Drive: End-to-End Autonomous Driving via Intention-Aware Physical Latent World ModelYupeng Zheng, Pengxuan Yang, Zebin Xing, Qichao Zhang et al.ICCV 2025 · 16 citations
- MP3: A Unified Model To Map, Perceive, Predict and PlanSergio Casas, Abbas Sadat, Raquel UrtasunCVPR 2021
- S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Model with Spatio-Temporal Visual RepresentationYichen Xie, Runsheng Xu, Tong He, Jyh-Jing Hwang et al.CVPR 2025
- Navigation-Guided Sparse Scene Representation for End-to-End Autonomous DrivingPeidong Li, Dixiao CuiICLR 2025
