Masked Spatio-Temporal Structure Prediction for Self-supervised Learning on Point Cloud Videos
Zhiqiang Shen, Xiaoxiao Sheng, Hehe Fan, Longguang Wang, Yulan Guo, Qiong Liu, Hao Wen, Xi Zhou
摘要
Recently, the community has made tremendous progress in developing effective methods for point cloud video understanding that learn from massive amounts of labeled data. However, annotating point cloud videos is usually notoriously expensive. Moreover, training via one or only a few traditional tasks (e.g., classification) may be insufficient to learn subtle details of the spatio-temporal structure existing in point cloud videos. In this paper, we propose a Masked Spatio-Temporal Structure Prediction (MaST-Pre) method to capture the structure of point cloud videos without human annotations. MaST-Pre is based on spatio-temporal pointtube masking and consists of two self-supervised learning tasks. First, by reconstructing masked point tubes, our method is able to capture the appearance information of point cloud videos. Second, to learn motion, we propose a temporal cardinality difference prediction task that estimates the change in the number of points within a point tube. In this way, MaST-Pre is forced to model the spatial and temporal structure in point cloud videos. Extensive experiments on MSRAction-3D, NTU-RGBD, NvGesture, and SHREC'17 demonstrate the effectiveness of the proposed method. The code is available at https://github. com/JohnsonSign/MaST-Pre .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- A Unified Framework for Human-centric Point Cloud Video UnderstandingYiteng Xu, Kecheng Ye, Xiao Han, Yiming Ren 等CVPR 2024 · 被引用 4 次
- DeSPITE: Exploring Contrastive Deep Skeleton-Pointcloud-IMU-Text Embeddings for Advanced Point Cloud Human Activity UnderstandingThomas Kreutz, Max Mühlhäuser, Alejandro Sánchez GuineaICCV 2025 · 被引用 1 次
- Adapting Pre-trained 3D Models for Point Cloud Video Understanding via Cross-frame Spatio-temporal PerceptionBaixuan Lv, Yaohua Zha, Tao Dai, Xue Yuerong 等CVPR 2025
- Mamba4D: Efficient 4D Point Cloud Video Understanding with Disentangled Spatial-Temporal State Space ModelsJiuming Liu, Jinru Han, Lihao Liu, Angelica I. Avilés-Rivero 等CVPR 2025
- PvNeXt: Rethinking Network Design and Temporal Motion for Point Cloud Video RecognitionJie Wang, Tingfa Xu, Lihe Ding, Xinjie Zhang 等ICLR 2025
它引用的顶会 Paper26
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
相关 Paper
- PointCMP: Contrastive Mask Prediction for Self-supervised Learning on Point Cloud VideosZhiqiang Shen, Xiaoxiao Sheng, Longguang Wang, Yulan Guo 等CVPR 2023
- Point 4D Transformer Networks for Spatio-Temporal Modeling in Point Cloud VideosHehe Fan, Yi Yang, Mohan S. KankanhalliCVPR 2021
- PSTNet: Point Spatio-Temporal Convolution on Point Cloud SequencesHehe Fan, Xin Yu, Yuhang Ding, Yi Yang 等ICLR 2021 · 被引用 148 次
- GeoMAE: Masked Geometric Target Prediction for Self-supervised Point Cloud Pre-TrainingXiaoyu Tian, Haoxi Ran, Yue Wang, Hang ZhaoCVPR 2023
- Self-Supervised Pillar Motion Learning for Autonomous DrivingChenxu Luo, Xiaodong Yang, Alan L. YuilleCVPR 2021
