PVT-SSD: Single-Stage 3D Object Detector with Point-Voxel Transformer
Honghui Yang, Wenxiao Wang, Minghao Chen, Binbin Lin, Tong He, Hua Chen, Xiaofei He, Wanli Ouyang
摘要
Recent Transformer-based 3D object detectors learn point cloud features either from point-or voxel-based representations. However, the former requires time-consuming sampling while the latter introduces quantization errors. In this paper, we present a novel Point-Voxel Transformer for single-stage 3D detection (PVT-SSD) that takes advantage of these two representations. Specifically, we first use voxel-based sparse convolutions for efficient feature encoding. Then, we propose a Point-Voxel Transformer (PVT) module that obtains long-range contexts in a cheap manner from voxels while attaining accurate positions from points. The key to associating the two different representations is our introduced input-dependent Query Initialization module, which could efficiently generate reference points and content queries. Then, PVT adaptively fuses long-range contextual and local geometric information around reference points into content queries. Further, to quickly find the neighboring points of reference points, we design the Virtual Range Image module, which generalizes the native range image to multi-sensor and multi-frame. The experiments on several autonomous driving benchmarks verify the effectiveness and efficiency of the proposed method. Code will be available at https:// github.com/ Nightmare-n/PVT-SSD.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- LION: Linear Group RNN for 3D Object Detection in Point CloudsZhe Liu, Jinghua Hou, Xinyu Wang, Xiaoqing Ye 等NeurIPS 2024 · 被引用 84 次
- UniPAD: A Universal Pre-Training Paradigm for Autonomous DrivingHonghui Yang, Sha Zhang, Di Huang, Xiaoyang Wu 等CVPR 2024 · 被引用 31 次
- LitePT: Lighter Yet Stronger Point TransformerYuanwen Yue, Damien Robert, Jianyuan Wang, Sunghwan Hong 等CVPR 2026 · 被引用 25 次
- AS-Det: Active Sampling for Adaptive 3D Object Detection in Point CloudsZiheng Ding, Xiaze Zhang, Qi Jing, Ying Cheng 等AAAI 2025 · 被引用 2 次
- WinMamba: Multi-Scale Shifted Windows in State Space Model for 3D Object DetectionLonghui Zheng, Qiming Xia, Xiaolu Chen, Zhaoliang Liu 等AAAI 2026 · 被引用 2 次
它引用的顶会 Paper52
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 被引用 1,467 次
- Voxel R-CNN: Towards High Performance Voxel-based 3D Object DetectionJiajun Deng, Shaoshuai Shi, Peiwei Li, Wengang Zhou 等AAAI 2021 · 被引用 1,128 次
- STD: Sparse-to-Dense 3D Object Detector for Point CloudZetong Yang, Yanan Sun, Shu Liu, Xiaoyong Shen 等ICCV 2019 · 被引用 840 次
相关 Paper
- Voxel Transformer for 3D Object DetectionJiageng Mao, Yujing Xue, Minzhe Niu, Haoyue Bai 等ICCV 2021 · 被引用 535 次
- MsSVT: Mixed-scale Sparse Voxel Transformer for 3D Object Detection on Point CloudsShaocong Dong, Lihe Ding, Haiyang Wang, Tingfa Xu 等NeurIPS 2022 · 被引用 37 次
- Voxel Set Transformer: A Set-to-Set Approach to 3D Object Detection from Point CloudsChenhang He, Ruihuang Li, Shuai Li, Lei ZhangCVPR 2022 · 被引用 217 次
- Fast Point R-CNNYilun Chen, Shu Liu, Xiaoyong Shen, Jiaya JiaICCV 2019 · 被引用 440 次
- HVNet: Hybrid Voxel Network for LiDAR Based 3D Object DetectionMaosheng Ye, Shuangjie Xu, Tongyi CaoCVPR 2020
