Towards foundational LiDAR world models with efficient latent flow matching
Tianran Liu, Shengwen Zhao, Nicholas Rhinehart
Abstract
LiDAR-based world models offer more structured and geometry-aware representations than their image-based counterparts. However, existing LiDAR world models are narrowly trained; each model excels only in the domain for which it was built. Can we develop LiDAR world models that exhibit strong transferability across multiple domains? We conduct the first systematic domain transfer study across three demanding scenarios: (i) outdoor to indoor generalization, (ii) sparse-beam&dense-beam adaptation, and (iii) non-semantic to semantic transfer. Given different amounts of fine-tuning data, our experiments show that a single pre-trained model can achieve up to 11% absolute improvement (83% relative) over training from scratch and outperforms training from scratch in 30/36 of our comparisons. This transferability of dynamic learning significantly reduces the reliance on manually annotated data for semantic occupancy forecasting: our method exceed the previous semantic occupancy forecasting models with only 5% of the labeled training data required by prior models. We also observed inefficiencies of current LiDAR world models, mainly through their under-compression of LiDAR data and inefficient training objectives. To address this, we propose a latent conditional flow matching (CFM)-based frameworks that achieves state-of-the-art reconstruction accuracy using only half the training data and a compression ratio 6 times higher than that of prior methods. Our model achieves SOTA performance on future-trajectory-conditioned semantic occupancy forecasting while being 23x more computationally efficient (a 28x FPS speedup); and achieves SOTA performance on semantic occupancy forecasting while being 2x more computationally efficient (a 1.1x FPS speedup).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5364d186-2a3b-4c3d-9903-881455765c4cCited by top-tier papers5
- LiDARCrafter: Dynamic 4D World Modeling from LiDAR SequencesAlan Liang, Youquan Liu, Yu Yang, Dongyue Lu et al.AAAI 2026 · 12 citations
- Mean Flow Distillation: Robust and Stable Distillation for Flow Matching ModelsAn Zhao, Shengyuan Zhang, Zhongjian Sun, Yixiang Zhou et al.ICML 2026 · 2 citations
- Causal Flow Q-Learning for Robust Offline Reinforcement LearningMingxuan Li, Junzhe Zhang, Elias BareinboimICML 2026 · 1 citation
- UCM: Unified Modeling of Camera Control and Memory with Time-aware Positional Encoding Warping for World ModelsTianxing Xu, Zi-Xuan Wang, Guangyuan Wang, Li Hu et al.SIGGRAPH 2026
- FM-Steer: Enhance Generalist Policies with Value-Guided Cascaded DenoisingHaoming Song, Delin Qu, Yuanqi Yao, Qizhi Chen et al.CVPR 2026
Builds on23
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
Related papers
- SG-LDM: Semantic-Guided LiDAR Generation via Latent-Aligned DiffusionZhengkang Xiang, Zizhao Li, Amir Khodabandeh, Kourosh KhoshelhamICCV 2025
- Temporal Overlapping Prediction: A Self-Supervised Pre-Training Method for LiDAR Moving Object SegmentationZiliang Miao, Runjian Chen, Yixi Cai, Buwei He et al.ICCV 2025 · 1 citation
- DIO: Decomposable Implicit 4D Occupancy-Flow World ModelChristopher Diehl, Quinlan Sykora, Ben Agro, Thomas Gilles et al.CVPR 2025
- Semi-supervised 3D Semantic Scene Completion with 2D Vision Foundation Model GuidanceDuc-Hai Pham, Duc Dung Nguyen, Anh Pham, Tuan Ho et al.AAAI 2025 · 6 citations
- Towards Realistic Scene Generation with LiDAR Diffusion ModelsHaoxi Ran, Vitor Guizilini, Yue WangCVPR 2024 · 26 citations
