Unifying Voxel-based Representation with Transformer for 3D Object Detection
Yanwei Li, Yilun Chen, Xiaojuan Qi, Zeming Li, Jian Sun, Jiaya Jia
摘要
In this work, we present a unified framework for multi-modality 3D object detection, named UVTR. The proposed method aims to unify multi-modality representations in the voxel space for accurate and robust single-or cross-modality 3D detection. To this end, the modality-specific space is first designed to represent different inputs in the voxel feature space. Different from previous work, our approach preserves the voxel space without height compression to alleviate semantic ambiguity and enable spatial connections. To make full use of the inputs from different sensors, the cross-modality interaction is then proposed, including knowledge transfer and modality fusion. In this way, geometry-aware expressions in point clouds and context-rich features in images are well utilized for better performance and robustness. The transformer decoder is applied to efficiently sample features from the unified space with learnable positions, which facilitates object-level interactions. In general, UVTR presents an early attempt to represent different modalities in a unified framework. It surpasses previous work in single-or multi-modality entries. The proposed method achieves leading performance in the nuScenes test set for both object detection and the following object tracking task. Code is made publicly available at https://github.com/dvlab-research/UVTR . 36th Conference on Neural Information Processing Systems (NeurIPS 2022).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper92
- Exploring Object-Centric Temporal Modeling for Efficient Multi-View 3D Object DetectionShihao Wang, Yingfei Liu, Tiancai Wang, Ying Li 等ICCV 2023 · 被引用 399 次
- SparseBEV: High-Performance Sparse 3D Object Detection from Multi-Camera VideosHaisong Liu, Yao Teng, Tao Lu, Haiguang Wang 等ICCV 2023 · 被引用 204 次
- FB-BEV: BEV Representation from Forward-Backward View TransformationsZhiqi Li, Zhiding Yu, Wenhai Wang, Anima Anandkumar 等ICCV 2023 · 被引用 144 次
- Cross Modal Transformer: Towards Fast and Robust 3D Object DetectionJunjie Yan, Yingfei Liu, Jianjian Sun, Fan Jia 等ICCV 2023 · 被引用 143 次
- Voxel Mamba: Group-Free State Space Models for Point Cloud based 3D Object DetectionGuowen Zhang, Lue Fan, Chenhang He, Zhen Lei 等NeurIPS 2024 · 被引用 137 次
它引用的顶会 Paper21
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 被引用 1,467 次
- Voxel R-CNN: Towards High Performance Voxel-based 3D Object DetectionJiajun Deng, Shaoshuai Shi, Peiwei Li, Wengang Zhou 等AAAI 2021 · 被引用 1,128 次
- STD: Sparse-to-Dense 3D Object Detector for Point CloudZetong Yang, Yanan Sun, Shu Liu, Xiaoyong Shen 等ICCV 2019 · 被引用 840 次
- A Comprehensive Overhaul of Feature DistillationByeongho Heo, Jeesoo Kim, Sangdoo Yun, Hyojin Park 等ICCV 2019 · 被引用 727 次
相关 Paper
- Uni3DETR: Unified 3D Detection TransformerZhenyu Wang, Ya-Li Li, Xi Chen, Hengshuang Zhao 等NeurIPS 2023 · 被引用 65 次
- UniTR: A Unified and Efficient Multi-Modal Transformer for Bird's-Eye-View RepresentationHaiyang Wang, Hao Tang, Shaoshuai Shi, Aoxue Li 等ICCV 2023 · 被引用 106 次
- Voxel Field Fusion for 3D Object DetectionYanwei Li, Xiaojuan Qi, Yukang Chen, Liwei Wang 等CVPR 2022 · 被引用 112 次
- Viewpoint Equivariance for Multi-View 3D Object DetectionDian Chen, Jie Li, Vitor Guizilini, Rares Ambrus 等CVPR 2023
- Every View Counts: Cross-View Consistency in 3D Object Detection with Hybrid-Cylindrical-Spherical VoxelizationQi Chen, Lin Sun, Ernest Cheung, Alan L. YuilleNeurIPS 2020 · 被引用 124 次
