Language Driven Occupancy Prediction
Zhu Yu, Bowen Pang, Lizhe Liu, Runmin Zhang, Qiang Li, Si-Yuan Cao, Maochun Luo, Mingxia Chen, Sheng Yang, Hui-Liang Shen
摘要
We introduce LOcc, an effective and generalizable framework for open-vocabulary occupancy (OVO) prediction. Previous approaches typically supervise the networks through coarse voxel-to-text correspondences via image features as intermediates or noisy and sparse correspondences from voxel-based model-view projections. To alleviate the inaccurate supervision, we propose a semantic transitive labeling pipeline to generate dense and fine-grained 3D language occupancy ground truth. Our pipeline presents a feasible way to dig into the valuable semantic information of images, transferring text labels from images to LiDAR point clouds and ultimately to voxels, to establish precise voxel-to-text correspondences. By replacing the original prediction head of supervised occupancy models with a geometry head for binary occupancy states and a language head for language features, LOcc effectively uses the generated language ground truth to guide the learning of 3D language volume. Through extensive experiments, we demonstrate that our transitive semantic labeling pipeline can produce more accurate pseudo-labeled ground truth, diminishing labor-intensive human annotations. Additionally, we validate LOcc across various architectures, where all models consistently outperform state-of-the-art zero-shot occupancy prediction approaches on the Occ3D-nuScenes dataset.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- GaussianFlowOcc: Sparse and Weakly Supervised Occupancy Estimation using Gaussian Splatting and Temporal FlowSimon Boeder, Fabian Gigengack, Benjamin RisseICCV 2025 · 被引用 28 次
- Large Depth Completion Model from Sparse ObservationsZhu Yu, zhengyi zhao, Runmin Zhang, Lingteng Qiu 等ICLR 2026 · 被引用 8 次
- Monocular Open Vocabulary Occupancy Prediction for Indoor ScenesChangqing Zhou, Yueru Luo, Han Zhang, Zeyu Jiang 等CVPR 2026 · 被引用 7 次
- See through the Dark: Learning Illumination-affined Representations for Nighttime Occupancy PredictionYuan Wu, Zhiqiang Yan, Yigong Zhang, Xiang Li 等NeurIPS 2025 · 被引用 7 次
- ShelfOcc: Native 3D Supervision beyond LiDAR for Vision-Based Occupancy EstimationSimon Boeder, Fabian Gigengack, Simon Roesler, Holger Caesar 等CVPR 2026 · 被引用 7 次
它引用的顶会 Paper34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel 等ICCV 2019 · 被引用 2,345 次
相关 Paper
- POP-3D: Open-Vocabulary 3D Occupancy Prediction from ImagesAntonín Vobecký, Oriane Siméoni, David Hurych, Spyridon Gidaris 等NeurIPS 2023 · 被引用 67 次
- AGO: Adaptive Grounding for Open World 3D Occupancy PredictionPeizheng Li, Shuxiao Ding, You Zhou, Qingwen Zhang 等ICCV 2025 · 被引用 4 次
- SurroundOcc: Multi-Camera 3D Occupancy Prediction for Autonomous DrivingYi Wei, Linqing Zhao, Wenzhao Zheng, Zheng Zhu 等ICCV 2023 · 被引用 380 次
- AutoOcc: Automatic Open-Ended Semantic Occupancy Annotation via Vision-Language Guided Gaussian SplattingXiaoyu Zhou, Jingqi Wang, Yongtao Wang, Yufei Wei 等ICCV 2025 · 被引用 1 次
- FusionOcc: Multi-Modal Fusion for 3D Occupancy PredictionShuo Zhang, Yupeng Zhai, Jilin Mei, Yu HuACM MM 2024 · 被引用 5 次
