X-Drive: Cross-modality Consistent Multi-Sensor Data Synthesis for Driving Scenarios
Yichen Xie, Chenfeng Xu, Chensheng Peng, Shuqi Zhao, Nhat Ho, Alexander T. Pham, Mingyu Ding, Masayoshi Tomizuka, Wei Zhan
摘要
Recent advancements have exploited diffusion models for the synthesis of either LiDAR point clouds or camera image data in driving scenarios. Despite their success in modeling single-modality data marginal distribution, there is an under-exploration in the mutual reliance between different modalities to describe complex driving scenes. To fill in this gap, we propose a novel framework, X-DRIVE, to model the joint distribution of point clouds and multi-view images via a dual-branch latent diffusion model architecture. Considering the distinct geometrical spaces of the two modalities, X-DRIVE conditions the synthesis of each modality on the corresponding local regions from the other modality, ensuring better alignment and realism. To further handle the spatial ambiguity during denoising, we design the cross-modality condition module based on epipolar lines to adaptively learn the cross-modality local correspondence. Besides, X-DRIVE allows for controllable generation through multi-level input conditions, including text, bounding box, image, and point clouds. Extensive results demonstrate the high-fidelity synthetic results of X-DRIVE for both point clouds and multi-view images, adhering to input conditions while ensuring reliable cross-modality consistency. Our code will be made publicly available at https://github.com/yichen928/X-Drive.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- X-Scene: Large-Scale Driving Scene Generation with High Fidelity and Flexible ControllabilityYu Yang, Alan Liang, Jianbiao Mei, Yukai Ma 等NeurIPS 2025 · 被引用 22 次
- LiDARCrafter: Dynamic 4D World Modeling from LiDAR SequencesAlan Liang, Youquan Liu, Yu Yang, Dongyue Lu 等AAAI 2026 · 被引用 12 次
- RLGF: Reinforcement Learning with Geometric Feedback for Autonomous Driving Video GenerationTianyi Yan, Wencheng Han, Xia Zhou, Xueyang Zhang 等NeurIPS 2025 · 被引用 9 次
- RAYNOVA: Scale-Temporal Autoregressive World Modeling in Ray SpaceYichen Xie, Chensheng Peng, Mazen Abdelfattah, Yihan Hu 等CVPR 2026 · 被引用 5 次
- Sensor2Sensor: Cross-Embodiment Sensor Conversion for Autonomous DrivingJiahao Wang, Bo Sun, Yijing Bai, Vincent Casser 等CVPR 2026 · 被引用 2 次
它引用的顶会 Paper33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- Towards Realistic Scene Generation with LiDAR Diffusion ModelsHaoxi Ran, Vitor Guizilini, Yue WangCVPR 2024 · 被引用 26 次
- Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal ConsistencyXiangyu Guo, Zhanqian Wu, Kaixin Xiong, Ziyang Xu 等NeurIPS 2025 · 被引用 24 次
- SG-LDM: Semantic-Guided LiDAR Generation via Latent-Aligned DiffusionZhengkang Xiang, Zizhao Li, Amir Khodabandeh, Kourosh KhoshelhamICCV 2025
- Autoscape: Geometry-Consistent Long-Horizon Scene GenerationJiacheng Chen, Ziyu Jiang, Mingfu Liang, Bingbing Zhuang 等ICCV 2025
- OmniGen: Unified Multimodal Sensor Generation for Autonomous DrivingTao Tang, Enhui Ma, Xia Zhou, Letian Wang 等ACM MM 2025 · 被引用 1 次
