X4D-SceneFormer: Enhanced Scene Understanding on 4D Point Cloud Videos through Cross-Modal Knowledge Transfer
Linglin Jing, Ying Xue, Xu Yan, Chaoda Zheng, Dong Wang, Ruimao Zhang, Zhigang Wang, Hui Fang, Bin Zhao, Zhen Li
Abstract
The field of 4D point cloud understanding is rapidly developing with the goal of analyzing dynamic 3D point cloud sequences. However, it remains a challenging task due to the sparsity and lack of texture in point clouds. Moreover, the irregularity of point cloud poses a difficulty in aligning temporal information within video sequences. To address these issues, we propose a novel cross-modal knowledge transfer framework, called X4D-SceneFormer. This framework enhances 4D-Scene understanding by transferring texture priors from RGB sequences using a Transformer architecture with temporal relationship mining. Specifically, the framework is designed with a dual-branch architecture, consisting of an 4D point cloud transformer and a Gradient-aware Image Transformer (GIT). The GIT combines visual texture and temporal correlation features to offer rich semantics and dynamics for better point cloud representation. During training, we employ multiple knowledge transfer techniques, including temporal consistency losses and masked self-attention, to strengthen the knowledge transfer between modalities. This leads to enhanced performance during inference using singlemodal 4D point cloud inputs. Extensive experiments demonstrate the superior performance of our framework on various 4D point cloud video understanding tasks, including action recognition, action segmentation and semantic segmentation. The results achieve 1st places, i.e., 85.3% (+7.9%) accuracy and 47.3% (+5.0%) mIoU for 4D action segmentation and semantic segmentation, on the HOI4D challenge 1 , outperforming previous state-of-the-art by a large margin. We release the code at https://github.com/jinglinglingling/X4D
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1cd0f2a0-1b78-42e9-ba5b-6d1e2ac72627Builds on14
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- CrossPoint: Self-Supervised Cross-Modal Contrastive Learning for 3D Point Cloud UnderstandingMohamed Afham, Isuru Dissanayake, Dinithi Dissanayake, Amaya Dharmasiri et al.CVPR 2022 · 286 citations
- MeteorNet: Deep Learning on Dynamic 3D Point Cloud SequencesXingyu Liu, Mengyuan Yan, Jeannette BohgICCV 2019 · 225 citations
- Point-to-Voxel Knowledge Distillation for LiDAR Semantic SegmentationYuenan Hou, Xinge Zhu, Yuexin Ma, Chen Change Loy et al.CVPR 2022 · 185 citations
Related papers
- Mamba4D: Efficient 4D Point Cloud Video Understanding with Disentangled Spatial-Temporal State Space ModelsJiuming Liu, Jinru Han, Lihao Liu, Angelica I. Avilés-Rivero et al.CVPR 2025
- Point 4D Transformer Networks for Spatio-Temporal Modeling in Point Cloud VideosHehe Fan, Yi Yang, Mohan S. KankanhalliCVPR 2021
- SpSequenceNet: Semantic Segmentation Network on 4D Point CloudsHanyu Shi, Guosheng Lin, Hao Wang, Tzu-Yi Hung et al.CVPR 2020
- Adapting Pre-trained 3D Models for Point Cloud Video Understanding via Cross-frame Spatio-temporal PerceptionBaixuan Lv, Yaohua Zha, Tao Dai, Xue Yuerong et al.CVPR 2025
- PSTNet: Point Spatio-Temporal Convolution on Point Cloud SequencesHehe Fan, Xin Yu, Yuhang Ding, Yi Yang et al.ICLR 2021 · 148 citations
