OmniGen: Unified Multimodal Sensor Generation for Autonomous Driving
Tao Tang, Enhui Ma, Xia Zhou, Letian Wang, Tianyi Yan, Xueyang Zhang, Kun Zhan, Peng Jia, Xianpeng Lang, Jia-Wang Bian, Kaicheng Yu, Xiaodan Liang
Abstract
Autonomous driving has seen remarkable advancements, largely driven by extensive real-world data collection. However, acquiring diverse and corner-case data remains costly and inefficient. Generative models have emerged as a promising solution by synthesizing realistic sensor data. However, existing approaches primarily focus on single-modality generation, leading to inefficiencies and misalignment in multimodal sensor data. To address these challenges, we propose OminiGen, which generates aligned multimodal sensor data in a unified framework. Our approach leverages a shared Bird's Eye View (BEV) space to unify multimodal features and designs a novel generalizable multimodal reconstruction method, UAE, to jointly decode LiDAR and multi-view camera data. UAE achieves multimodal sensor decoding through volume rendering, enabling accurate and flexible reconstruction. Furthermore, we incorporate a Diffusion Transformer (DiT) with a ControlNet branch to enable controllable multimodal sensor generation. Our comprehensive experiments demonstrate that OminiGen achieves desired performances in unified multimodal sensor data generation with multimodal consistency and flexible sensor adjustments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 716996ed-a14c-4f68-bfd5-c9eabfe60dd9Cited by top-tier papers2
- Sensor2Sensor: Cross-Embodiment Sensor Conversion for Autonomous DrivingJiahao Wang, Bo Sun, Yijing Bai, Vincent Casser et al.CVPR 2026 · 2 citations
- CorrectAD: A Self-Correcting Agentic System to Improve End-to-end Planning in Autonomous DrivingEnhui Ma, Lijun Zhou, Tao Tang, Jiahuan Zhang et al.AAAI 2026
Builds on46
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- X-Drive: Cross-modality Consistent Multi-Sensor Data Synthesis for Driving ScenariosYichen Xie, Chenfeng Xu, Chensheng Peng, Shuqi Zhao et al.ICLR 2025
- Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal ConsistencyXiangyu Guo, Zhanqian Wu, Kaixin Xiong, Ziyang Xu et al.NeurIPS 2025 · 24 citations
- SurfelGAN: Synthesizing Realistic Sensor Data for Autonomous DrivingZhenpei Yang, Yuning Chai, Dragomir Anguelov, Yin Zhou et al.CVPR 2020
- StreetCrafter: Street View Synthesis with Controllable Video Diffusion ModelsYunzhi Yan, Zhen Xu, Haotong Lin, Haian Jin et al.CVPR 2025
- DiffE2E: Rethinking End-to-End Driving with a Hybrid Diffusion-Regression-Classification PolicyRui Zhao, Yuze Fan, Ziguo Chen, Fei Gao et al.NeurIPS 2025 · 7 citations
