SSLFusion: Scale and Space Aligned Latent Fusion Model for Multimodal 3D Object Detection
Bonan Ding, Jin Xie, Jing Nie, Jiale Cao
摘要
Multimodal 3D object detection based on deep neural networks has indeed made significant progress. However, it still faces challenges due to the misalignment of scale and spatial information between features extracted from 2D images and those derived from 3D point clouds. Existing methods usually aggregate multimodal features at a single stage. However, leveraging multi-stage cross-modal features is crucial for detecting objects of various scales. Therefore, these methods often struggle to integrate features across different scales and modalities effectively, thereby restricting the accuracy of detection. Additionally, the time-consuming Query-Key-Value-based (QKV-based) cross-attention operations often utilized in existing methods aid in reasoning the location and existence of objects by capturing non-local contexts. However, this approach tends to increase computational complexity. To address these challenges, we present SSLFusion, a novel Scale & Space Aligned Latent Fusion Model, consisting of a scale-aligned fusion strategy (SAF), a 3D-to-2D space alignment module (SAM), and a latent cross-modal fusion module (LFM). SAF mitigates scale misalignment between modalities by aggregating features from both images and point clouds across multiple levels. SAM is designed to reduce the inter-modal gap between features from images and point clouds by incorporating 3D coordinate information into 2D image features. Additionally, LFM captures crossmodal non-local contexts in the latent space without utilizing the QKV-based attention operations, thus mitigating computational complexity. Experiments on the KITTI and DENSE datasets demonstrate that our SSLFusion outperforms stateof-the-art methods. Our approach obtains an absolute gain of 2.15% in 3D AP, compared with the state-of-art method GraphAlign on the moderate level of the KITTI test set.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper9
- Voxel R-CNN: Towards High Performance Voxel-based 3D Object DetectionJiajun Deng, Shaoshuai Shi, Peiwei Li, Wengang Zhou 等AAAI 2021 · 被引用 1,128 次
- Fog Simulation on Real LiDAR Point Clouds for 3D Object Detection in Adverse WeatherMartin Hahner, Christos Sakaridis, Dengxin Dai, Luc Van GoolICCV 2021 · 被引用 210 次
- CAT-Det: Contrastively Augmented Transformer for Multimodal 3D Object DetectionYanan Zhang, Jiaxin Chen, Di HuangCVPR 2022 · 被引用 138 次
- GraphAlign: Enhancing Accurate Feature Alignment by Graph matching for Multi-Modal 3D Object DetectionZiying Song, Haiyue Wei, Lin Bai, Lei Yang 等ICCV 2023 · 被引用 73 次
- PV-RCNN: Point-Voxel Feature Set Abstraction for 3D Object DetectionShaoshuai Shi, Chaoxu Guo, Li Jiang, Zhe Wang 等CVPR 2020
相关 Paper
- DLFusion: Painting-Depth Augmenting-LiDAR for Multimodal Fusion 3D Object DetectionJunyin Wang, Chenghu Du, Hui Li, Shengwu XiongACM MM 2023 · 被引用 3 次
- LoGoNet: Towards Accurate 3D Object Detection with Local-to-Global Cross- Modal FusionXin Li, Tao Ma, Yuenan Hou, Botian Shi 等CVPR 2023
- Sparse Fuse Dense: Towards High Quality 3D Detection with Depth CompletionXiaopei Wu, Liang Peng, Honghui Yang, Liang Xie 等CVPR 2022 · 被引用 248 次
- Accurate Monocular 3D Object Detection via Color-Embedded 3D Reconstruction for Autonomous DrivingXinzhu Ma, Zhihui Wang, Haojie Li, Pengbo Zhang 等ICCV 2019 · 被引用 339 次
- ObjectFusion: Multi-modal 3D Object Detection with Object-Centric FusionQi Cai, Yingwei Pan, Ting Yao, Chong-Wah Ngo 等ICCV 2023 · 被引用 71 次
