TopNet: Transformer-Efficient Occupancy Prediction Network for Octree-Structured Point Cloud Geometry Compression
Xinjie Wang, Yifan Zhang, Ting Liu, Xinpu Liu, Ke Xu, Jianwei Wan, Yulan Guo, Hanyun Wang
摘要
Efficient Point Cloud Geometry Compression (PCGC) with a lower bits per point (BPP) and higher peak signalto-noise ratio (PSNR) is essential for the transportation of large-scale 3D data. Although octree-based entropy models can reduce BPP without introducing geometry distortion, existing CNN-based models struggle with limited receptive fields to capture long-range dependencies, while Transformer-built architectures always neglect fine-grained details due to their reliance on global selfattention. In this paper, we propose a Transformerefficient occupancy prediction Network, termed TopNet, to overcome these challenges by developing several novel components: Locally-enhanced Context Encoding (LeCE) for enhancing the translation-invariance of the octree nodes, Adaptive-Length Sliding Window Attention (AL-SWA) for capturing both global and local dependencies while adaptively adjusting attention weights based on the input window length, Spatial-Gated-enhanced Channel Mixer (SG-CM) for efficient feature aggregation from ancestors and siblings, and Latent-guided Node Occupancy Predictor (LNOP) for improving prediction accuracy of spatially adjacent octree nodes. Comprehensive experiments across both indoor and outdoor point cloud datasets demonstrate that our TopNet achieves state-ofthe-art performance with fewer parameters, further advancing the reduction-efficiency boundaries of PCGC. The code is available at https : / / github . com / xinjiewang1995/TopNet.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- AnyPcc: Compressing Any Point Cloud with a Single Universal ModelKangli Wang, Qianxi Yi, Yuqi Ye, Shihao Li 等CVPR 2026 · 被引用 4 次
- ELiC: Efficient LiDAR Geometry Compression via Cross-Bit-depth Feature Propagation and Bag-of-EncodersJunsik Kim, Gun Bang, Soowoong KimCVPR 2026 · 被引用 1 次
它引用的顶会 Paper18
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel 等ICCV 2019 · 被引用 2,345 次
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao 等CVPR 2022 · 被引用 2,138 次
- CMT: Convolutional Neural Networks Meet Vision TransformersJianyuan Guo, Kai Han, Han Wu, Yehui Tang 等CVPR 2022 · 被引用 839 次
- Not All Points Are Equal: Learning Highly Efficient Point-based Detectors for 3D LiDAR Point CloudsYifan Zhang, Qingyong Hu, Guoquan Xu, Yanxin Ma 等CVPR 2022 · 被引用 376 次
相关 Paper
- OctFormer: Efficient Octree-Based Transformer for Point Cloud Compression with Local EnhancementMingyue Cui, Junhua Long, Mingjian Feng, Boyang Li 等AAAI 2023 · 被引用 56 次
- Efficient Hierarchical Entropy Model for Learned Point Cloud CompressionRui Song, Chunyang Fu, Shan Liu, Ge LiCVPR 2023
- OctAttention: Octree-Based Large-Scale Contexts Model for Point Cloud CompressionChunyang Fu, Ge Li, Rui Song, Wei Gao 等AAAI 2022 · 被引用 191 次
- VoxelContext-Net: An Octree Based Framework for Point Cloud CompressionZizheng Que, Guo Lu, Dong XuCVPR 2021
- OctFormer: Octree-based Transformers for 3D Point CloudsPeng-Shuai WangSIGGRAPH 2023 · 被引用 123 次
