Height-Fidelity Dense Global Fusion for Multi-Modal 3D Object Detection
Hanshi Wang, Jin Gao, Weiming Hu, Zhipeng Zhang
摘要
We present the first work demonstrating that a pure Mamba block can achieve efficient Dense Global Fusion, meanwhile guaranteeing top performance for cameraLiDAR multi-modal 3D object detection. Our motivation stems from the observation that existing fusion strategies are constrained by their inability to simultaneously achieve efficiency, long-range modeling, and retaining complete scene information. Inspired by recent advances in statespace models (SSMs) [8] and linear attention [35], [43], we leverage their linear complexity and long-range modeling capabilities to address these challenges. However, this is non-trivial since our experiments reveal that simply adopting efficient linear-complexity methods does not necessarily yield improvements and may even degrade performance. We attribute this degradation to the loss of height information during multi-modal alignment, leading to deviations in sequence order. To resolve this, we propose height-fidelity LiDAR encoding that preserves precise height information through voxel compression in continuous space, thereby enhancing camera-LiDAR alignment. Subsequently, we introduce the Hybrid Mamba Block, which leverages the enriched height-informed features to conduct local and global contextual learning. By integrating these components, our method achieves state-of-the-art performance with the top-tire NDS score of 75.0 on the nuScenes [2] validation benchmark, even surpassing methods that utilize highresolution inputs. Meanwhile, our method maintains efficiency, achieving faster inference speed than most recent state-of-the-art methods. Code is available at https://github.com/AutoLab-SAI-SJTU/MambaFusion
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Hi-Gaussian: Hierarchical Gaussians Under Normalized Spherical Projection for Single-View 3D ReconstructionBinjian Xie, Pengju Zhang, Hao Wei, Yihong WuICCV 2025 · 被引用 2 次
- MI-TRQR: Mutual Information-Based Temporal Redundancy Quantification and Reduction for Energy-Efficient Spiking Neural NetworksDengfeng Xue, Wenjuan Li, Yifan Lu, Chunfeng Yuan 等NeurIPS 2025
它引用的顶会 Paper30
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu 等NeurIPS 2024 · 被引用 3,199 次
- CCNet: Criss-Cross Attention for Semantic SegmentationZilong Huang, Xinggang Wang, Lichao Huang, Chang Huang 等ICCV 2019 · 被引用 2,972 次
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang 等ICML 2024 · 被引用 1,725 次
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 被引用 1,407 次
相关 Paper
- UniMamba: Unified Spatial-Channel Representation Learning with Group-Efficient Mamba for LiDAR-based 3D Object DetectionXin Jin, Haisheng Su, Kai Liu, Cong Ma 等CVPR 2025
- MSMDFusion: Fusing LiDAR and Camera at Multiple Scales with Multi-Depth Seeds for 3D Object DetectionYang Jiao, Zequn Jie, Shaoxiang Chen, Jingjing Chen 等CVPR 2023
- SparseFusion: Fusing Multi-Modal Sparse Representations for Multi-Sensor 3D Object DetectionYichen Xie, Chenfeng Xu, Marie-Julie Rakotosaona, Patrick Rim 等ICCV 2023 · 被引用 134 次
- GAFusion: Adaptive Fusing LiDAR and Camera with Multiple Guidance for 3D Object DetectionXiaotian Li, Baojie Fan, Jiandong Tian, Huijie FanCVPR 2024
- CRAFT: Camera-Radar 3D Object Detection with Spatio-Contextual Fusion TransformerYoungseok Kim, Sanmin Kim, Jun Won Choi, Dongsuk KumAAAI 2023 · 被引用 145 次
