H2GFormer: Horizontal-to-Global Voxel Transformer for 3D Semantic Scene Completion
Yu Wang, Chao Tong
Abstract
3D Semantic Scene Completion (SSC) has emerged as a novel task in vision-based holistic 3D scene understanding. Its objective is to densely predict the occupancy and category of each voxel in a 3D scene based on input from either LiDAR or images. Currently, many transformer-based semantic scene completion frameworks employ simple yet popular Cross-Attention and Self-Attention mechanisms to integrate and infer dense geometric and semantic information of voxels. However, they overlook the distinctions among voxels in the scene, especially in outdoor scenarios where the horizontal direction contains more variations. And voxels located at object boundaries and within the interior of objects exhibit varying levels of positional significance. To address this issue, we propose a transformer-based SSC framework called H2GFormer that incorporates a horizontal-to-global approach. This framework takes into full consideration the variations of voxels in the horizontal direction and the characteristics of voxels on object boundaries. We introduce a horizontal window-to-global attention (W2G) module that effectively fuses semantic information by first diffusing it horizontally from reliably visible voxels and then propagating the semantic understanding to global voxels, ensuring a more reliable fusion of semantic-aware features. Moreover, an Internal-External Position Awareness Loss (IoE-PALoss) is utilized during network training to emphasize the critical positions within the transition regions between objects. The experiments conducted on the SemanticKITTI dataset demonstrate that H2GFormer exhibits superior performance in both geometric and semantic completion tasks. Our code is available on https://github.com/Ryanwy1/H2GFormer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4257fae5-2b84-4ee5-b49d-9f80959281a2Cited by top-tier papers19
- Context and Geometry Aware Voxel Transformer for Semantic Scene CompletionZhu Yu, Runmin Zhang, Jiacheng Ying, Junchen Yu et al.NeurIPS 2024 · 73 citations
- LowRankOcc: Tensor Decomposition and Low-Rank Recovery for Vision-Based 3D Semantic Occupancy PredictionLinqing Zhao, Xiuwei Xu, Ziwei Wang, Yunpeng Zhang et al.CVPR 2024 · 14 citations
- VLScene: Vision-Language Guidance Distillation for Camera-Based 3D Semantic Scene CompletionMeng Wang, Huilong Pi, Ruihui Li, Yunchuan Qin et al.AAAI 2025 · 11 citations
- Skip Mamba Diffusion for Monocular 3D Semantic Scene CompletionLi Liang, Naveed Akhtar, Jordan Vice, Xiangrui Kong et al.AAAI 2025 · 10 citations
- OccAny: Generalized Unconstrained Urban 3D OccupancyAnh-Quan Cao, Tuan-Hung VuCVPR 2026 · 6 citations
Builds on18
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel et al.ICCV 2019 · 2,345 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- SurroundOcc: Multi-Camera 3D Occupancy Prediction for Autonomous DrivingYi Wei, Linqing Zhao, Wenzhao Zheng, Zheng Zhu et al.ICCV 2023 · 380 citations
- Sparse Single Sweep LiDAR Point Cloud Segmentation via Learning Contextual Shape Priors from Scene CompletionXu Yan, Jiantao Gao, Jie Li, Ruimao Zhang et al.AAAI 2021 · 365 citations
Related papers
- VoxFormer: Sparse Voxel Transformer for Camera-Based 3D Semantic Scene CompletionYiming Li, Zhiding Yu, Christopher B. Choy, Chaowei Xiao et al.CVPR 2023
- OccFormer: Dual-path Transformer for Vision-based 3D Semantic Occupancy PredictionYunpeng Zhang, Zheng Zhu, Dalong DuICCV 2023 · 354 citations
- SGFormer: Semantic-Geometry Fusion Transformer for Multi-modal 3D Panoptic SegmentationHongqi Yu, Sixian Chan, Xiaolong Zhou, Xiaoqin ZhangAAAI 2025 · 3 citations
- IPFormer: Visual 3D Panoptic Scene Completion with Context-Adaptive Instance ProposalsMarkus Gross, Aya Fahmy, Danit Niwattananan, Dominik Muhle et al.NeurIPS 2025 · 3 citations
- Point Cloud Semantic Scene Completion with Prototype-Guided TransformerChenghao Fang, Jianqing Liang, Jiye Liang, Zijin Du et al.AAAI 2026
