Cross Modal Transformer: Towards Fast and Robust 3D Object Detection
Junjie Yan, Yingfei Liu, Jianjian Sun, Fan Jia, Shuailin Li, Tiancai Wang, Xiangyu Zhang
摘要
In this paper, we propose a robust 3D detector, named Cross Modal Transformer (CMT), for end-to-end 3D multi-modal detection. Without explicit view transformation, CMT takes the image and point clouds tokens as inputs and directly outputs accurate 3D bounding boxes. The spatial alignment of multi-modal tokens is performed by encoding the 3D points into multi-modal features. The core design of CMT is quite simple while its performance is impressive. It achieves 74.1% NDS (state-of-the-art with single model) on nuScenes test set while maintaining faster inference speed. Moreover, CMT has a strong robustness even if the LiDAR is missing. Code is released at https://github.com/junjie18/CMT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper33
- TUMTraf V2X Cooperative Perception DatasetWalter Zimmer, Gerhard Arya Wardana, Suren Sritharan, Xingcheng Zhou 等CVPR 2024 · 被引用 76 次
- HGSFusion: Radar-Camera Fusion with Hybrid Generation and Synchronization for 3D Object DetectionZijian Gu, Jianwei Ma, Yan Huang, Honghao Wei 等AAAI 2025 · 被引用 26 次
- CRKD: Enhanced Camera-Radar Object Detection with Cross-Modality Knowledge DistillationLingjun Zhao, Jingyu Song, Katherine A. SkinnerCVPR 2024 · 被引用 21 次
- QE-BEV: Query Evolution for Bird's Eye View Object Detection in Varied ContextsJiawei Yao, Yingxin Lai, Hongrui Kou, Tong Wu 等ACM MM 2024 · 被引用 18 次
- BEVSpread: Spread Voxel Pooling for Bird's-Eye-View Representation in Vision-Based Roadside 3D Object DetectionWenjie Wang, Yehao Lu, Guangcong Zheng, Shuigen Zhan 等CVPR 2024 · 被引用 17 次
它引用的顶会 Paper27
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang 等ICLR 2022 · 被引用 1,218 次
- BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object DetectionYinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang 等AAAI 2023 · 被引用 954 次
- TrackFormer: Multi-Object Tracking with TransformersTim Meinhardt, Alexander Kirillov, Laura Leal-Taixé, Christoph FeichtenhoferCVPR 2022 · 被引用 927 次
- DN-DETR: Accelerate DETR Training by Introducing Query DeNoisingFeng Li, Hao Zhang, Shilong Liu, Jian Guo 等CVPR 2022 · 被引用 879 次
相关 Paper
- Unifying Voxel-based Representation with Transformer for 3D Object DetectionYanwei Li, Yilun Chen, Xiaojuan Qi, Zeming Li 等NeurIPS 2022 · 被引用 401 次
- Bridged Transformer for Vision and Point Cloud 3D Object DetectionYikai Wang, TengQi Ye, Lele Cao, Wenbing Huang 等CVPR 2022 · 被引用 55 次
- CAT-Det: Contrastively Augmented Transformer for Multimodal 3D Object DetectionYanan Zhang, Jiaxin Chen, Di HuangCVPR 2022 · 被引用 138 次
- PointAugmenting: Cross-Modal Augmentation for 3D Object DetectionChunwei Wang, Chao Ma, Ming Zhu, Xiaokang YangCVPR 2021
- EVT: Efficient View Transformation for Multi-Modal 3D Object DetectionYongjin Lee, Hyeon Mun Jeong, Yurim Jeon, Sanghyun KimICCV 2025 · 被引用 5 次
