OcTr: Octree-Based Transformer for 3D Object Detection
Chao Zhou, Yanan Zhang, Jiaxin Chen, Di Huang
摘要
A key challenge for LiDAR-based 3D object detection is to capture sufficient features from large scale 3D scenes especially for distant or/and occluded objects. Albeit recent efforts made by Transformers with the long sequence modeling capability, they fail to properly balance the accuracy and efficiency, suffering from inadequate receptive fields or coarse-grained holistic correlations. In this paper, we propose an Octree-based Transformer, named OcTr, to address this issue. It first constructs a dynamic octree on the hierarchical feature pyramid through conducting self-attention on the top level and then recursively propagates to the level below restricted by the octants, which captures rich global context in a coarse-to-fine manner while maintaining the computational complexity under control. Furthermore, for enhanced foreground perception, we propose a hybrid positional embedding, composed of the semantic-aware positional embedding and attention mask, to fully exploit semantic and geometry clues. Extensive experiments are conducted on the Waymo Open Dataset and KITTI Dataset, and OcTr reaches newly state-of-the-art results.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- HEDNet: A Hierarchical Encoder-Decoder Network for 3D Object Detection in Point CloudsGang Zhang, Junnan Chen, Guohuan Gao, Jianmin Li 等NeurIPS 2023 · 被引用 95 次
- LION: Linear Group RNN for 3D Object Detection in Point CloudsZhe Liu, Jinghua Hou, Xinyu Wang, Xiaoqing Ye 等NeurIPS 2024 · 被引用 84 次
- IS-Fusion: Instance-Scene Collaborative Fusion for Multimodal 3D Object DetectionJunbo Yin, Jianbing Shen, Runnan Chen, Wei Li 等CVPR 2024 · 被引用 73 次
- OctreeOcc: Efficient and Multi-Granularity Occupancy Prediction Using Octree QueriesYuhang Lu, Xinge Zhu, Tai Wang, Yuexin MaNeurIPS 2024 · 被引用 70 次
- SA-BEV: Generating Semantic-Aware Bird's-Eye-View Feature for Multi-view 3D Object DetectionJinqing Zhang, Yanan Zhang, Qingjie Liu, Yunhong WangICCV 2023 · 被引用 41 次
它引用的顶会 Paper36
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 被引用 1,467 次
- CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped WindowsXiaoyi Dong, Jianmin Bao, Dongdong Chen, Weiming Zhang 等CVPR 2022 · 被引用 1,207 次
- Voxel R-CNN: Towards High Performance Voxel-based 3D Object DetectionJiajun Deng, Shaoshuai Shi, Peiwei Li, Wengang Zhou 等AAAI 2021 · 被引用 1,128 次
相关 Paper
- SEFormer: Structure Embedding Transformer for 3D Object DetectionXiaoyu Feng, Heming Du, Hehe Fan, Yueqi Duan 等AAAI 2023 · 被引用 15 次
- HOTFormerLoc: Hierarchical Octree Transformer for Versatile Lidar Place Recognition Across Ground and Aerial ViewsEthan Griffiths, Maryam Haghighat, Simon Denman, Clinton Fookes 等CVPR 2025
- Voxel Transformer for 3D Object DetectionJiageng Mao, Yujing Xue, Minzhe Niu, Haoyue Bai 等ICCV 2021 · 被引用 535 次
- OctFormer: Octree-based Transformers for 3D Point CloudsPeng-Shuai WangSIGGRAPH 2023 · 被引用 123 次
- MsSVT: Mixed-scale Sparse Voxel Transformer for 3D Object Detection on Point CloudsShaocong Dong, Lihe Ding, Haiyang Wang, Tingfa Xu 等NeurIPS 2022 · 被引用 37 次
