Monocular Open Vocabulary Occupancy Prediction for Indoor Scenes
Changqing Zhou, Yueru Luo, Han Zhang, Zeyu Jiang, Changhao Chen
摘要
Open-vocabulary 3D occupancy is vital for embodied agents, which need to understand complex indoor environments where semantic categories are abundant and evolve beyond fixed taxonomies. While recent work has explored open-vocabulary occupancy in outdoor driving scenarios, such methods transfer poorly indoors, where geometry is denser, layouts are more intricate, and semantics are far more fine-grained. To address these challenges, we adopt a geometry-only supervision paradigm that uses only binary occupancy labels (occupied vs free). Our framework builds upon 3D Language-Embedded Gaussians, which serve as a unified intermediate representation coupling fine-grained 3D geometry with a language-aligned semantic embedding. On the geometry side, we find that existing Gaussian-to-Occupancy operators fail to converge under such weak supervision, and we introduce an opacity-aware, Poisson-based approach that stabilizes volumetric aggregation. On the semantic side, direct alignment between rendered features and open-vocabulary segmentation features suffers from feature mixing; we therefore propose a Progressive Temperature Decay schedule that gradually sharpens opacities during splatting, strengthening Gaussian-language alignment. On Occ-ScanNet, our framework achieves 59.50 IoU and 21.05 mIoU in the open-vocabulary setting, surpassing all existing occupancy methods in IoU and outperforming prior open-vocabulary approaches by a large margin in mIoU. Code will be released at https://github.com/JuIvyy/LegoOcc.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao 等NeurIPS 2024 · 被引用 2,305 次
- SurroundOcc: Multi-Camera 3D Occupancy Prediction for Autonomous DrivingYi Wei, Linqing Zhao, Wenzhao Zheng, Zheng Zhu 等ICCV 2023 · 被引用 380 次
- OccFormer: Dual-path Transformer for Vision-based 3D Semantic Occupancy PredictionYunpeng Zhang, Zheng Zhu, Dalong DuICCV 2023 · 被引用 354 次
相关 Paper
- EmbodiedOcc: Embodied 3D Occupancy Prediction for Vision-Based Online Scene UnderstandingYuqi Wu, Wenzhao Zheng, Sicheng Zuo, Yuanhui Huang 等ICCV 2025 · 被引用 4 次
- AutoOcc: Automatic Open-Ended Semantic Occupancy Annotation via Vision-Language Guided Gaussian SplattingXiaoyu Zhou, Jingqi Wang, Yongtao Wang, Yufei Wei 等ICCV 2025 · 被引用 1 次
- Language Driven Occupancy PredictionZhu Yu, Bowen Pang, Lizhe Liu, Runmin Zhang 等ICCV 2025 · 被引用 1 次
- GaussTR: Foundation Model-Aligned Gaussian Transformer for Self-Supervised 3D Spatial UnderstandingHaoyi Jiang, Liu Liu, Tianheng Cheng, Xinjie Wang 等CVPR 2025
- Dr. Splat: Directly Referring 3D Gaussian Splatting via Direct Language Embedding RegistrationKim Jun-Seong, GeonU Kim, Kim Yu-Ji, Yu-Chiang Frank Wang 等CVPR 2025
