Sensor2Sensor: Cross-Embodiment Sensor Conversion for Autonomous Driving
Jiahao Wang, Bo Sun, Yijing Bai, Vincent Casser, Songyou Peng, Zehao Zhu, Meng-Li Shih, Xander Masotto, Shih-Yang Su, Kanaad V. Parvate, Tiancheng Ge, Linn Bieske
Abstract
Robust training and validation of Autonomous Driving Systems (ADS) require massive, diverse datasets. Proprietary data collected by Autonomous Vehicle (AV) fleets, while high-fidelity, are limited in scale, diversity of sensor configurations, as well as geographic and long-tail-behavioral coverage. In contrast, in-the-wild data from sources like dashcams offers immense scale and diversity, capturing critical long-tail scenarios and novel environments. However, this unstructured, in-the-wild video data is incompatible with ADS expecting structured, multi-modal sensor inputs for validation and training. To bridge this data gap, we propose Sensor2Sensor, a novel generative model- † Work done during an internship at Waymo.
ing paradigm that translates in-the-wild monocular dashcam videos into a high-fidelity, multi-modal sensor suite (AV logs) comprising multi-view camera images and LiDAR point clouds. A core challenge is the lack of paired training data. We address this by converting real AV logs into dashcam-style videos via 4D Gaussian Splatting (4DGS) reconstruction and novel-view rendering. Sensor2Sensor then utilizes a diffusion architecture to perform the generative conversion. We perform comprehensive quantitative evaluations on the fidelity and realism of the generated sensor data. We demonstrate Sensor2Sensor's practical utility by converting challenging in-the-wild internet and dashcam footage into realistic, multi-modal data formats, further unlocking vast external data sources for AV development.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b6830fd8-0d26-4cad-9ad2-e85014b33d3bBuilds on26
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- 4D Gaussian Splatting for Real-Time Dynamic Scene RenderingGuanjun Wu, Taoran Yi, Jiemin Fang, Lingxi Xie et al.CVPR 2024 · 513 citations
- Genie: Generative Interactive EnvironmentsJake Bruce, Michael D. Dennis, Ashley Edwards, Jack Parker-Holder et al.ICML 2024 · 513 citations
Related papers
- SurfelGAN: Synthesizing Realistic Sensor Data for Autonomous DrivingZhenpei Yang, Yuning Chai, Dragomir Anguelov, Yin Zhou et al.CVPR 2020
- Unraveling the Effects of Synthetic Data on End-to-End Autonomous DrivingJunhao Ge, Zuhong Liu, Longteng Fan, Yifan Jiang et al.ICCV 2025 · 2 citations
- WorldSplat: Gaussian-Centric Feed-Forward 4D Scene Generation for Autonomous DrivingZiyue Zhu, Zhanqian Wu, Zhenxin Zhu, Lijun Zhou et al.ICLR 2026 · 11 citations
- StreetCrafter: Street View Synthesis with Controllable Video Diffusion ModelsYunzhi Yan, Zhen Xu, Haotong Lin, Haian Jin et al.CVPR 2025
- OmniGen: Unified Multimodal Sensor Generation for Autonomous DrivingTao Tang, Enhui Ma, Xia Zhou, Letian Wang et al.ACM MM 2025 · 1 citation
