PTNET: A Proposal-Centric Transformer Net-Work for 3D Object Detection
Jianping Zhong, Zhaobo Qi, Kaiwen Duan, Xinyan Liu, Beichen Zhang, Weigang Zhang, Qingming Huang
Abstract
3D object detection using LiDAR point cloud data is critical for autonomous driving systems. However, recent two-stage detectors still struggle to deliver satisfactory performance primarily due to inadequate proposal quality, which stems from significant geometric detail degradation in generated proposal features caused by high sparsity and uneven distribution of point clouds, as well as a complete failure to exploit surrounding contextual cues during independent proposal refinement, losing complementary details from adjacent proposals. To this end, we propose a Proposal-centric Transformer Network (PTN), which includes a Hierarchical Attentive Feature Alignment (HAFA) and a Collaborative Proposal Refinement Module (CPRM). More concretely, HAFA employs a dual-stream architecture to extract multi-granularity proposal representations, including coarse-grained multi-scale voxel features and fine-grained coordinate point features to enhance proposals' object geometric representation ability. CPRM first generates hybrid object queries for all objects and then establishes contextual-aware interactions through the 3D parameter-guided deformable attention mechanism to effectively aggregate spatial location cues and category-specific information across proposals that are spatially adjacent and semantically correlated. Extensive experiments on the large-scale Waymo and KITTI benchmarks demonstrate the superiority of PTN. The code is available at https://github.com/ZhongJianPing1/ ptnet.git .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on25
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Voxel R-CNN: Towards High Performance Voxel-based 3D Object DetectionJiajun Deng, Shaoshuai Shi, Peiwei Li, Wengang Zhou et al.AAAI 2021 · 1,128 citations
- TransFusion: Robust LiDAR-Camera Fusion for 3D Object Detection with TransformersXuyang Bai, Zeyu Hu, Xinge Zhu, Qingqiu Huang et al.CVPR 2022 · 794 citations
- Not All Points Are Equal: Learning Highly Efficient Point-based Detectors for 3D LiDAR Point CloudsYifan Zhang, Qingyong Hu, Guoquan Xu, Yanxin Ma et al.CVPR 2022 · 376 citations
- Fully Sparse 3D Object DetectionLue Fan, Feng Wang, Naiyan Wang, Zhaoxiang ZhangNeurIPS 2022 · 168 citations
Related papers
- Improving 3D Object Detection with Channel-wise TransformerHualian Sheng, Sijia Cai, Yuan Liu, Bing Deng et al.ICCV 2021 · 293 citations
- 3D Object Detection With PointformerXuran Pan, Zhuofan Xia, Shiji Song, Li Erran Li et al.CVPR 2021
- Point Density-Aware Voxels for LiDAR 3D Object DetectionJordan S. K. Hu, Tianshu Kuai, Steven L. WaslanderCVPR 2022
- GeoFormer: Geometry Point Encoder for 3D Object Detection with Graph-Based TransformerXin Jin, Haisheng Su, Cong Ma, Kai Liu et al.ICCV 2025 · 2 citations
- PC-RGNN: Point Cloud Completion and Graph Neural Network for 3D Object DetectionYanan Zhang, Di Huang, Yunhong WangAAAI 2021 · 109 citations
