PSTNet: Point Spatio-Temporal Convolution on Point Cloud Sequences
Hehe Fan, Xin Yu, Yuhang Ding, Yi Yang, Mohan S. Kankanhalli
Abstract
Point cloud sequences are irregular and unordered in the spatial dimension while exhibiting regularities and order in the temporal dimension. Therefore, existing grid based convolutions for conventional video processing cannot be directly applied to spatio-temporal modeling of raw point cloud sequences. In this paper, we propose a point spatio-temporal (PST) convolution to achieve informative representations of point cloud sequences. The proposed PST convolution first disentangles space and time in point cloud sequences. Then, a spatial convolution is employed to capture the local structure of points in the 3D space, and a temporal convolution is used to model the dynamics of the spatial regions along the time dimension. Furthermore, we incorporate the proposed PST convolution into a deep network, namely PSTNet, to extract features of point cloud sequences in a hierarchical manner. Extensive experiments on widely-used 3D action recognition and 4D semantic segmentation datasets demonstrate the effectiveness of PSTNet to model point cloud sequences.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d5ca6288-aea8-44f2-a89a-723503367da7Cited by top-tier papers35
- HOI4D: A 4D Egocentric Dataset for Category-Level Human-Object InteractionYunze Liu, Yun Liu, Che Jiang, Kangbo Lyu et al.CVPR 2022 · 126 citations
- From Chaos Comes Order: Ordering Event Representations for Object Recognition and DetectionNikola Zubic, Daniel Gehrig, Mathias Gehrig, Davide ScaramuzzaICCV 2023 · 71 citations
- Self-Supervised Global-Local Structure Modeling for Point Cloud Domain Adaptation with Reliable Voted Pseudo LabelsHehe Fan, Xiaojun Chang, Wanyue Zhang, Yi Cheng et al.CVPR 2022 · 61 citations
- GIFS: Neural Implicit Function for General Shape RepresentationJianglong Ye, Yuntao Chen, Naiyan Wang, Xiaolong WangCVPR 2022 · 54 citations
- Clustering based Point Cloud Representation Learning for 3D AnalysisTuo Feng, Wenguan Wang, Xiaohan Wang, Yi Yang et al.ICCV 2023 · 53 citations
Builds on9
- KPConv: Flexible and Deformable Convolution for Point CloudsHugues Thomas, Charles R. Qi, Jean-Emmanuel Deschaud, Beatriz Marcotegui et al.ICCV 2019 · 3,193 citations
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 1,467 citations
- Occupancy Flow: 4D Reconstruction by Learning Particle DynamicsMichael Niemeyer, Lars M. Mescheder, Michael Oechsle, Andreas GeigerICCV 2019 · 314 citations
- MeteorNet: Deep Learning on Dynamic 3D Point Cloud SequencesXingyu Liu, Mengyuan Yan, Jeannette BohgICCV 2019 · 225 citations
- A Universal Representation Transformer Layer for Few-Shot Image ClassificationLu Liu, William L. Hamilton, Guodong Long, Jing Jiang et al.ICLR 2021 · 143 citations
Related papers
- Point 4D Transformer Networks for Spatio-Temporal Modeling in Point Cloud VideosHehe Fan, Yi Yang, Mohan S. KankanhalliCVPR 2021
- SpSequenceNet: Semantic Segmentation Network on 4D Point CloudsHanyu Shi, Guosheng Lin, Hao Wang, Tzu-Yi Hung et al.CVPR 2020
- Mamba4D: Efficient 4D Point Cloud Video Understanding with Disentangled Spatial-Temporal State Space ModelsJiuming Liu, Jinru Han, Lihao Liu, Angelica I. Avilés-Rivero et al.CVPR 2025
- UST-SSM: Unified Spatio-Temporal State Space Models for Point Cloud Video ModelingPeiming Li, Ziyi Wang, Yulin Yuan, Hong Liu et al.ICCV 2025 · 3 citations
- Adapting Pre-trained 3D Models for Point Cloud Video Understanding via Cross-frame Spatio-temporal PerceptionBaixuan Lv, Yaohua Zha, Tao Dai, Xue Yuerong et al.CVPR 2025
