SEFormer: Structure Embedding Transformer for 3D Object Detection
Xiaoyu Feng, Heming Du, Hehe Fan, Yueqi Duan, Yongpan Liu
Abstract
Effectively preserving and encoding structure features from objects in irregular and sparse LiDAR points is a crucial challenge to 3D object detection on the point cloud. Recently, Transformer has demonstrated promising performance on many 2D and even 3D vision tasks. Compared with the fixed and rigid convolution kernels, the self-attention mechanism in Transformer can adaptively exclude the unrelated or noisy points and is thus suitable for preserving the local spatial structure in the irregular LiDAR point cloud. However, Transformer only performs a simple sum on the point features, based on the self-attention mechanism, and all the points share the same transformation for value. A such isotropic operation cannot capture the direction-distance-oriented local structure, which is essential for 3D object detection. In this work, we propose a Structure-Embedding transFormer (SEFormer), which can not only preserve the local structure as a traditional Transformer but also have the ability to encode the local structure. Compared to the self-attention mechanism in traditional Transformer, SEFormer learns different feature transformations for value points based on the relative directions and distances to the query point. Then we propose a SEFormer-based network for high-performance 3D object detection. Extensive experiments show that the proposed architecture can achieve SOTA results on the Waymo Open Dataset, one of the most significant 3D detection benchmarks for autonomous driving. Specifically, SEFormer achieves 79.02% mAP, which is 1.2% higher than existing works. https://github.com/tdzdog/SEFormer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 66639921-9a60-4080-b73d-ffe41481a2a8Cited by top-tier papers6
- SPGroup3D: Superpoint Grouping Network for Indoor 3D Object DetectionYun Zhu, Le Hui, Yaqi Shen, Jin XieAAAI 2024 · 24 citations
- SwiftPillars: High-Efficiency Pillar Encoder for Lidar-Based 3D DetectionXin Jin, Kai Liu, Cong Ma, Ruining Yang et al.AAAI 2024 · 12 citations
- How Many Tokens Do 3D Point Cloud Transformer Architectures Really Need?Tuan Anh Tran, Duy M. H. Nguyen, Hoai-Chau Tran, Michael Barz et al.NeurIPS 2025 · 5 citations
- PointListNet: Deep Learning on 3D Point ListsHehe Fan, Linchao Zhu, Yi Yang, Mohan S. KankanhalliCVPR 2023
- Object-Goal Visual Navigation via Effective Exploration of Relations Among Historical Navigation StatesHeming Du, Lincheng Li, Zi Huang, Xin YuCVPR 2023
Builds on28
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 1,467 citations
- STD: Sparse-to-Dense 3D Object Detector for Point CloudZetong Yang, Yanan Sun, Shu Liu, Xiaoyong Shen et al.ICCV 2019 · 840 citations
- Meta Faster R-CNN: Towards Accurate Few-Shot Object Detection with Attentive Feature AlignmentGuangxing Han, Shiyuan Huang, Jiawei Ma, Yicheng He et al.AAAI 2022 · 227 citations
- Behind the Curtain: Learning Occluded Shapes for 3D Object DetectionQiangeng Xu, Yiqi Zhong, Ulrich NeumannAAAI 2022 · 188 citations
- Every View Counts: Cross-View Consistency in 3D Object Detection with Hybrid-Cylindrical-Spherical VoxelizationQi Chen, Lin Sun, Ernest Cheung, Alan L. YuilleNeurIPS 2020 · 124 citations
Related papers
- GeoFormer: Geometry Point Encoder for 3D Object Detection with Graph-Based TransformerXin Jin, Haisheng Su, Cong Ma, Kai Liu et al.ICCV 2025 · 2 citations
- Embracing Single Stride 3D Object Detector with Sparse TransformerLue Fan, Ziqi Pang, Tianyuan Zhang, Yu-Xiong Wang et al.CVPR 2022
- 3D Object Detection With PointformerXuran Pan, Zhuofan Xia, Shiji Song, Li Erran Li et al.CVPR 2021
- OcTr: Octree-Based Transformer for 3D Object DetectionChao Zhou, Yanan Zhang, Jiaxin Chen, Di HuangCVPR 2023
- Point Density-Aware Voxels for LiDAR 3D Object DetectionJordan S. K. Hu, Tianshu Kuai, Steven L. WaslanderCVPR 2022
