Cross Modal Transformer: Towards Fast and Robust 3D Object Detection
Junjie Yan, Yingfei Liu, Jianjian Sun, Fan Jia, Shuailin Li, Tiancai Wang, Xiangyu Zhang
Abstract
In this paper, we propose a robust 3D detector, named Cross Modal Transformer (CMT), for end-to-end 3D multi-modal detection. Without explicit view transformation, CMT takes the image and point clouds tokens as inputs and directly outputs accurate 3D bounding boxes. The spatial alignment of multi-modal tokens is performed by encoding the 3D points into multi-modal features. The core design of CMT is quite simple while its performance is impressive. It achieves 74.1% NDS (state-of-the-art with single model) on nuScenes test set while maintaining faster inference speed. Moreover, CMT has a strong robustness even if the LiDAR is missing. Code is released at https://github.com/junjie18/CMT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 703e53cb-ae29-4d3f-af42-37203f1faf08Cited by top-tier papers33
- TUMTraf V2X Cooperative Perception DatasetWalter Zimmer, Gerhard Arya Wardana, Suren Sritharan, Xingcheng Zhou et al.CVPR 2024 · 76 citations
- HGSFusion: Radar-Camera Fusion with Hybrid Generation and Synchronization for 3D Object DetectionZijian Gu, Jianwei Ma, Yan Huang, Honghao Wei et al.AAAI 2025 · 26 citations
- CRKD: Enhanced Camera-Radar Object Detection with Cross-Modality Knowledge DistillationLingjun Zhao, Jingyu Song, Katherine A. SkinnerCVPR 2024 · 21 citations
- QE-BEV: Query Evolution for Bird's Eye View Object Detection in Varied ContextsJiawei Yao, Yingxin Lai, Hongrui Kou, Tong Wu et al.ACM MM 2024 · 18 citations
- BEVSpread: Spread Voxel Pooling for Bird's-Eye-View Representation in Vision-Based Roadside 3D Object DetectionWenjie Wang, Yehao Lu, Guangcong Zheng, Shuigen Zhan et al.CVPR 2024 · 17 citations
Builds on27
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang et al.ICLR 2022 · 1,218 citations
- BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object DetectionYinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang et al.AAAI 2023 · 954 citations
- TrackFormer: Multi-Object Tracking with TransformersTim Meinhardt, Alexander Kirillov, Laura Leal-Taixé, Christoph FeichtenhoferCVPR 2022 · 927 citations
- DN-DETR: Accelerate DETR Training by Introducing Query DeNoisingFeng Li, Hao Zhang, Shilong Liu, Jian Guo et al.CVPR 2022 · 879 citations
Related papers
- Unifying Voxel-based Representation with Transformer for 3D Object DetectionYanwei Li, Yilun Chen, Xiaojuan Qi, Zeming Li et al.NeurIPS 2022 · 401 citations
- Bridged Transformer for Vision and Point Cloud 3D Object DetectionYikai Wang, TengQi Ye, Lele Cao, Wenbing Huang et al.CVPR 2022 · 55 citations
- CAT-Det: Contrastively Augmented Transformer for Multimodal 3D Object DetectionYanan Zhang, Jiaxin Chen, Di HuangCVPR 2022 · 138 citations
- PointAugmenting: Cross-Modal Augmentation for 3D Object DetectionChunwei Wang, Chao Ma, Ming Zhu, Xiaokang YangCVPR 2021
- EVT: Efficient View Transformation for Multi-Modal 3D Object DetectionYongjin Lee, Hyeon Mun Jeong, Yurim Jeon, Sanghyun KimICCV 2025 · 5 citations
