SAM4D: Segment Anything in Camera and LiDAR Streams
Jianyun Xu, Song Wang, Ziqian Ni, Chunyong Hu, Sheng Yang, Jianke Zhu, Qiang Li
摘要
We present SAM4D, a multi-modal and temporal foundation model designed for promptable segmentation across camera and LiDAR streams. Unified Multi-modal Positional Encoding (UMPE) is introduced to align camera and LiDAR features in a shared 3D space, enabling seamless cross-modal prompting and interaction. Additionally, we propose Motion-aware Cross-modal Memory Attention (MCMA), which leverages ego-motion compensation to enhance temporal consistency and long-horizon feature retrieval, ensuring robust segmentation across dynamically changing autonomous driving scenes. To avoid annotation bottlenecks, we develop a multi-modal automated data engine that synergizes VFM-driven video masklets, spatiotemporal 4D reconstruction, and cross-modal masklet fusion. This framework generates camera-LiDAR aligned pseudo-labels at a speed orders of magnitude faster than human annotation while preserving VFM-derived semantic fidelity in point cloud representations. We conduct extensive experiments on the constructed Waymo-4DSeg, which demonstrate the powerful cross-modal segmentation ability and great potential in data annotation of proposed SAM4D.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- BEVFusion: A Simple and Robust LiDAR-Camera Fusion FrameworkTingting Liang, Hongwei Xie, Kaicheng Yu, Zhongyu Xia 等NeurIPS 2022 · 被引用 762 次
- Segment Anything in High QualityLei Ke, Mingqiao Ye, Martin Danelljan, Yifan Liu 等NeurIPS 2023 · 被引用 709 次
- Unifying Voxel-based Representation with Transformer for 3D Object DetectionYanwei Li, Yilun Chen, Xiaojuan Qi, Zeming Li 等NeurIPS 2022 · 被引用 401 次
相关 Paper
- LIFT: Learning 4D LiDAR Image Fusion Transformer for 3D Object DetectionYihan Zeng, Da Zhang, Chunwei Wang, Zhenwei Miao 等CVPR 2022 · 被引用 36 次
- M4-SAM: Multi-Modal Mixture-of-Experts with Memory-Augmented SAM for RGB-D Video Salient Object DetectionJiyuan Liu, Jia Lin, Xiaofei Zhou, Runmin Cong 等CVPR 2026
- MSeg3D: Multi-Modal 3D Semantic Segmentation for Autonomous DrivingJiale Li, Hang Dai, Hao Han, Yong DingCVPR 2023
- FusionSAM: Visual Multi-Modal Learning with Segment Anything ModelDaixun Li, Weiying Xie, Mingxiang Cao, Yunke Wang 等KDD 2025 · 被引用 2 次
- Beyond One Shot, Beyond One Perspective: Cross-View and Long-Horizon Distillation for Better LiDAR RepresentationsXiang Xu, Lingdong Kong, Song Wang, Chuanwei Zhou 等ICCV 2025 · 被引用 1 次
