LOMA: Language-assisted Semantic Occupancy Network via Triplane Mamba
Yubo Cui, Zhiheng Li, Jiaqiang Wang, Zheng Fang
Abstract
Vision-based 3D occupancy prediction has become a popular research task due to its versatility and affordability. Nowadays, conventional methods usually project the image-based vision features to 3D space and learn the geometric information through the attention mechanism, enabling the 3D semantic occupancy prediction. However, these works usually face two main challenges: 1) Limited geometric information. Due to the lack of geometric information in the image itself, it is challenging to directly predict 3D space information, especially in large-scale outdoor scenes. 2) Local restricted interaction. Due to the quadratic complexity of the attention mechanism, they often use modified local attention to fuse features, resulting in a restricted fusion. To address these problems, in this paper, we propose a language-assisted 3D semantic occupancy prediction network, named LOMA. In the proposed vision-language framework, we first introduce a VL-aware Scene Generator (VSG) module to generate the 3D language feature of the scene. By leveraging the vision-language model, this module provides implicit geometric knowledge and explicit semantic information from the language. Furthermore, we present a Tri-plane Fusion Mamba (TFM) block to efficiently fuse the 3D language feature and 3D vision feature. The proposed module not only fuses the two features with global modeling but also avoids too much computation costs. Experiments on the SemanticKITTI and SSCBench-KITTI360 datasets show that our algorithm achieves new state-of-the-art performances in both geometric and semantic completion tasks. Our code will be open soon.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f5941f7d-8b91-44c9-953d-860cb9bd8f3aCited by top-tier papers4
- MamTiff-CAD: Multi-Scale Latent Diffusion with Mamba+ for Complex Parametric SequenceLiyuan Deng, Yunpeng Bai, Yongkang Dai, Xiaoshui Huang et al.ICCV 2025 · 3 citations
- Unleashing Semantic and Geometric Priors for 3D Scene CompletionShiyuan Chen, Wei Sui, Bohao Zhang, Zeyd Boukhers et al.AAAI 2026 · 1 citation
- Towards 3D Object-Centric Feature Learning for Semantic Scene CompletionWeihua Wang, Yubo Cui, Xiangru Lin, Zhiheng Li et al.AAAI 2026
- SPSC: Sparse and Scalable Multi-Modal 3D Occupancy Prediction for Autonomous DrivingQingju Guo, Shuang Li, Binhui Xie, Jing Geng et al.AAAI 2026
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu et al.NeurIPS 2024 · 3,199 citations
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel et al.ICCV 2019 · 2,345 citations
Related papers
- OccMamba: Semantic Occupancy Prediction with State Space ModelsHeng Li, Yuenan Hou, Xiaohan Xing, Yuexin Ma et al.CVPR 2025
- AutoOcc: Automatic Open-Ended Semantic Occupancy Annotation via Vision-Language Guided Gaussian SplattingXiaoyu Zhou, Jingqi Wang, Yongtao Wang, Yufei Wei et al.ICCV 2025 · 1 citation
- VLScene: Vision-Language Guidance Distillation for Camera-Based 3D Semantic Scene CompletionMeng Wang, Huilong Pi, Ruihui Li, Yunchuan Qin et al.AAAI 2025 · 11 citations
- AGO: Adaptive Grounding for Open World 3D Occupancy PredictionPeizheng Li, Shuxiao Ding, You Zhou, Qingwen Zhang et al.ICCV 2025 · 4 citations
- Semi-supervised 3D Semantic Scene Completion with 2D Vision Foundation Model GuidanceDuc-Hai Pham, Duc Dung Nguyen, Anh Pham, Tuan Ho et al.AAAI 2025 · 6 citations
