How Do Images Align and Complement LiDAR? Towards a Harmonized Multi-modal 3D Panoptic Segmentation
Yining Pan, Qiongjie Cui, Xulei Yang, Na Zhao
摘要
LiDAR-based 3D panoptic segmentation often struggles with the inherent sparsity of data from LiDAR sensors, which makes it challenging to accurately recognize distant or small objects. Recently, a few studies have sought to overcome this challenge by integrating LiDAR inputs with camera images, leveraging the rich and dense texture information provided by the latter. While these approaches have shown promising results, they still face challenges, such as misalignment during data augmentation and the reliance on postprocessing steps. To address these issues, we propose Image-Assists-LiDAR (IAL), a novel multimodal 3D panoptic segmentation framework. In IAL, we first introduce a modality-synchronized data augmentation strategy, PieAug, to ensure alignment between LiDAR and image inputs from the start. Next, we adopt a transformer decoder to directly predict panoptic segmentation results. To effectively fuse LiDAR and image features into tokens for the decoder, we design a Geometricguided Token Fusion (GTF) module. Additionally, we leverage the complementary strengths of each modality as priors for query initialization through a Prior-based Query Generation (PQG) module, enhancing the decoder's ability to generate accurate instance masks. Our IAL framework achieves state-of-the-art performance compared to previous multi-modal 3D panoptic segmentation methods on two widely used benchmarks. Code and models are publicly available at https://github.com/IMPL-Lab/IAL.git .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Robust Multi-View Learning via Representation Fusion of Sample-Level Attention and Alignment of Simulated PerturbationJie Xu, Na Zhao, Gang Niu, Masashi Sugiyama 等ICCV 2025 · 被引用 6 次
- CCF: Complementary Collaborative Fusion for Domain Generalized Multi-Modal 3D Object DetectionYuchen Wu, Kun Wang, Yining Pan, Na ZhaoCVPR 2026 · 被引用 4 次
- Few-Shot Incremental 3D Object Detection in Dynamic Indoor EnvironmentsYun Zhu, Jianjun Qian, Jian Yang, Jin Xie 等CVPR 2026 · 被引用 2 次
- PanDA: Unsupervised Domain Adaptation for Multimodal 3D Panoptic Segmentation in Autonomous DrivingYining Pan, Shijie Li, Yuchen Wu, Xulei Yang 等CVPR 2026 · 被引用 1 次
- Graph Smoothing for Enhanced Local Geometry Learning in Point Cloud AnalysisShangbo Yuan, Jie Xu, Ping Hu, Xiaofeng Zhu 等AAAI 2026
它引用的顶会 Paper31
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel 等ICCV 2019 · 被引用 2,345 次
- TransFusion: Robust LiDAR-Camera Fusion for 3D Object Detection with TransformersXuyang Bai, Zeyu Hu, Xinge Zhu, Qingqiu Huang 等CVPR 2022 · 被引用 794 次
- BEVFusion: A Simple and Robust LiDAR-Camera Fusion FrameworkTingting Liang, Hongwei Xie, Kaicheng Yu, Zhongyu Xia 等NeurIPS 2022 · 被引用 762 次
- DeepFusion: Lidar-Camera Deep Fusion for Multi-Modal 3D Object DetectionYingwei Li, Adams Wei Yu, Tianjian Meng, Benjamin Caine 等CVPR 2022 · 被引用 508 次
相关 Paper
- MSeg3D: Multi-Modal 3D Semantic Segmentation for Autonomous DrivingJiale Li, Hang Dai, Hao Han, Yong DingCVPR 2023
- LiDAR-Camera Panoptic Segmentation via Geometry-Consistent and Semantic-Aware AlignmentZhiwei Zhang, Zhizhong Zhang, Qian Yu, Ran Yi 等ICCV 2023 · 被引用 26 次
- UniSeg: A Unified Multi-Modal LiDAR Segmentation Network and the OpenPCSeg CodebaseYouquan Liu, Runnan Chen, Xin Li, Lingdong Kong 等ICCV 2023 · 被引用 94 次
- CAT-Det: Contrastively Augmented Transformer for Multimodal 3D Object DetectionYanan Zhang, Jiaxin Chen, Di HuangCVPR 2022 · 被引用 138 次
- LIFT: Learning 4D LiDAR Image Fusion Transformer for 3D Object DetectionYihan Zeng, Da Zhang, Chunwei Wang, Zhenwei Miao 等CVPR 2022 · 被引用 36 次
