OneOcc: Semantic Occupancy Prediction for Legged Robots with a Single Panoramic Camera
Hao Shi, Ze Wang, Shangwei Guo, Mengfei Duan, Song Wang, Teng Chen, Kailun Yang, Lin Wang, Kaiwei Wang
Abstract
Robust 3D semantic occupancy is essential for legged and humanoid robots, yet most Semantic Scene Completion (SSC) systems are built for wheeled platforms with forward-facing sensors. We present , a vision-only panoramic SSC framework tailored to severe body jitter and continuity. OneOcc integrates four complementary modules: (i) , which jointly exploits the raw annular panorama and its equirectangular unfolding to preserve true continuity while enabling grid-aligned feature extraction and seam-aware context; (ii) , which reasons in Cartesian and polar/cylindrical voxel spaces to reduce discretization bias and better align with panoramic geometry, yielding sharper free/occupied boundaries; (iii) a lightweight decoder with fusion that dynamically routes multi-scale 3D features to specialized experts, improving long-range context and occlusion handling; and (iv) a plug-and-play module that learns feature-level motion correction from gait, stabilizing representations without extra sensors. We also release two panoramic occupancy benchmarks: (real quadruped, first-person ) and (CARLA human-ego with RGB/Depth/semantic-occupancy and standardized within-/cross-city splits). OneOcc sets new SOTA: on QuadOcc it exceeds strong vision baselines and even popular LiDAR methods, and on H3O it improves within-city by +3.83 mIoU and cross-city by +8.08. The modules are lightweight, enabling deployable full-surround semantic perception for legged and humanoid robots. Datasets and code will be released upon publication.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- TGSFormer: Scalable Temporal Gaussian Splatting for Embodied Semantic Scene CompletionRui Qian, Haozhi Cao, Tianchen Deng, TIANXIN HU et al.CVPR 2026 · 2 citations
- Gallant: Voxel Grid-based Humanoid Locomotion and Local-navigation across 3-D Constrained TerrainsQingwei Ben, Botian Xu, Kailin Li, Feiyu Jia et al.CVPR 2026
Builds on33
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel et al.ICCV 2019 · 2,345 citations
Related papers
- H2GFormer: Horizontal-to-Global Voxel Transformer for 3D Semantic Scene CompletionYu Wang, Chao TongAAAI 2024 · 33 citations
- OccuFly: A 3D Vision Benchmark for Semantic Scene Completion from the Aerial PerspectiveMarkus Gross, Sai B. Matha, Aya Fahmy, Rui Song et al.CVPR 2026 · 7 citations
- OccFormer: Dual-path Transformer for Vision-based 3D Semantic Occupancy PredictionYunpeng Zhang, Zheng Zhu, Dalong DuICCV 2023 · 354 citations
- EmbodiedOcc: Embodied 3D Occupancy Prediction for Vision-Based Online Scene UnderstandingYuqi Wu, Wenzhao Zheng, Sicheng Zuo, Yuanhui Huang et al.ICCV 2025 · 4 citations
- LowRankOcc: Tensor Decomposition and Low-Rank Recovery for Vision-Based 3D Semantic Occupancy PredictionLinqing Zhao, Xiuwei Xu, Ziwei Wang, Yunpeng Zhang et al.CVPR 2024 · 14 citations
