PvNeXt: Rethinking Network Design and Temporal Motion for Point Cloud Video Recognition
Jie Wang, Tingfa Xu, Lihe Ding, Xinjie Zhang, Long Bai, Jianan Li
Abstract
Point cloud video perception has become an essential task for the realm of 3D vision. Current 4D representation learning techniques typically engage in iterative processing coupled with dense query operations. Although effective in capturing temporal features, this approach leads to substantial computational redundancy. In this work, we propose a framework, named as PvNeXt, for effective yet efficient point cloud video recognition, via personalized one-shot query operation. Specially, PvNeXt consists of two key modules, the Motion Imitator and the Single-Step Motion Encoder. The former module, the Motion Imitator, is designed to capture the temporal dynamics inherent in sequences of point clouds, thus generating the virtual motion corresponding to each frame. The Single-Step Motion Encoder performs a one-step query operation, associating point cloud of each frame with its corresponding virtual motion frame, thereby extracting motion cues from point cloud sequences and capturing temporal dynamics across the entire sequence. Through the integration of these two modules, PvNeXt enables personalized one-shot queries for each frame, effectively eliminating the need for frame-specific looping and intensive query processes. Extensive experiments on multiple benchmarks demonstrate the effectiveness of our method. INTRODUCTION Point cloud videos serve as a pivotal character, offering a dynamic perspective into our environment, which is fundamental in the realms of robotics and AR systems. These sequences, which present movements within the physical domain, are crucial in delineating environmental transformations and facilitating interactions within said environments. This contrasts starkly with the limited descriptive capabilities of 2D images or static 3D point clouds. Therefore, enhancing the ability of point cloud video perception becomes a significant yet challenging task. However, 4D data representation learning presents vital challenges and remains a nascent field of inquiry. The amalgamation of 3D geometry and dynamic motion often leads to data redundancy within an exceedingly high-dimensional space, which heavily hinders the development of efficient spatio-temporal representations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d76b4ed6-8eab-4ba9-a831-64358faa2dceCited by top-tier papers1
Ask how each one uses itBuilds on17
- PointNeXt: Revisiting PointNet++ with Improved Training and Scaling StrategiesGuocheng Qian, Yuchen Li, Houwen Peng, Jinjie Mai et al.NeurIPS 2022 · 1,270 citations
- Channel-wise Topology Refinement Graph Convolution for Skeleton-Based Action RecognitionYuxin Chen, Ziqi Zhang, Chunfeng Yuan, Bing Li et al.ICCV 2021 · 871 citations
- Rethinking Network Design and Local Geometry in Point Cloud: A Simple Residual MLP FrameworkXu Ma, Can Qin, Haoxuan You, Haoxi Ran et al.ICLR 2022 · 841 citations
- MeteorNet: Deep Learning on Dynamic 3D Point Cloud SequencesXingyu Liu, Mengyuan Yan, Jeannette BohgICCV 2019 · 225 citations
- PSTNet: Point Spatio-Temporal Convolution on Point Cloud SequencesHehe Fan, Xin Yu, Yuhang Ding, Yi Yang et al.ICLR 2021 · 148 citations
Related papers
- Point 4D Transformer Networks for Spatio-Temporal Modeling in Point Cloud VideosHehe Fan, Yi Yang, Mohan S. KankanhalliCVPR 2021
- SpSequenceNet: Semantic Segmentation Network on 4D Point CloudsHanyu Shi, Guosheng Lin, Hao Wang, Tzu-Yi Hung et al.CVPR 2020
- Complete-to-Partial 4D Distillation for Self-Supervised Point Cloud Sequence Representation LearningZhuoyang Zhang, Yuhao Dong, Yunze Liu, Li YiCVPR 2023
- LeaF: Learning Frames for 4D Point Cloud Sequence UnderstandingYunze Liu, Junyu Chen, Zekai Zhang, Jingwei Huang et al.ICCV 2023 · 19 citations
- A Unified Framework for Human-centric Point Cloud Video UnderstandingYiteng Xu, Kecheng Ye, Xiao Han, Yiming Ren et al.CVPR 2024 · 4 citations
