Mamba4D: Efficient 4D Point Cloud Video Understanding with Disentangled Spatial-Temporal State Space Models
Jiuming Liu, Jinru Han, Lihao Liu, Angelica I. Avilés-Rivero, Chaokang Jiang, Zhe Liu, Hesheng Wang
摘要
Point cloud videos can faithfully capture real-world spatial geometries and temporal dynamics, which are essential for enabling intelligent agents to understand the dynamically changing world. However, designing an effective 4D backbone remains challenging, mainly due to the irregular and unordered distribution of points and temporal inconsistencies across frames. Also, recent transformerbased 4D backbones commonly suffer from large computational costs due to their quadratic complexity, particularly for long video sequences. To address these challenges, we propose a novel point cloud video understanding backbone purely based on the State Space Models (SSMs). Specifically, we first disentangle space and time in 4D video sequences and then establish the spatio-temporal correlation with the unified spatial-temporal Mamba blocks. The Intraframe Spatial Mamba module is developed to encode locally similar geometric structures within a certain temporal stride. Subsequently, locally correlated tokens are delivered to the Inter-frame Temporal Mamba module, which integrates long-term point features across the entire video with linear complexity. Our proposed Mamba4D achieves competitive performance on the MSR-Action3D action recognition (+10.4% accuracy), HOI4D action segmentation (+0.7 F1 Score), and Synthia4D semantic segmentation (+0.19 mIoU) datasets. Mamba4D also has a significant efficiency improvement, especially for long video sequences, with 87.5% GPU memory reduction and ×5.36 speed-up. Codes are released at https://github.com/IRMVLab/Mamba4D.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- SwiftVLA: Unlocking Spatiotemporal Dynamics for Lightweight VLA Models at Minimal OverheadChaojun Ni, Chen Cheng, Xiaofeng Wang, Zheng Zhu 等CVPR 2026 · 被引用 23 次
- Turboreg: Turboclique for Robust and Efficient Point Cloud RegistrationShaocheng Yan, Pengcheng Shi, Zhenjun Zhao, Kaixin Wang 等ICCV 2025 · 被引用 11 次
- 4DPChat: Towards Dynamic Point Cloud Understanding with Failure-Aware BootstrappingXindan Zhang, Weilong Yan, YUFEI SHI, Xuerui Qiu 等ICML 2026 · 被引用 6 次
- CloudMamba: Grouped Selective State Spaces for Point Cloud AnalysisKanglin Qu, Pan Gao, Qun Dai, Zhanzhi Ye 等AAAI 2026 · 被引用 2 次
- 4DSTR: Advancing Generative 4D Gaussians with Spatial-Temporal Rectification for High-Quality and Consistent 4D GenerationMengmeng Liu, Jiuming Liu, Yunpeng Zhang, Jiangtao Li 等AAAI 2026 · 被引用 2 次
它引用的顶会 Paper25
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu 等NeurIPS 2024 · 被引用 3,199 次
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang 等ICML 2024 · 被引用 1,725 次
- Combining Recurrent, Convolutional, and Continuous-time Models with Linear State Space LayersAlbert Gu, Isys Johnson, Karan Goel, Khaled Saab 等NeurIPS 2021 · 被引用 1,280 次
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 被引用 1,248 次
相关 Paper
- Pamba: Enhancing Global Interaction in Point Clouds via State Space ModelZhuoyuan Li, Yubo Ai, Jiahao Lu, Chuxin Wang 等AAAI 2025 · 被引用 12 次
- PointMamba: A Simple State Space Model for Point Cloud AnalysisDingkang Liang, Xin Zhou, Wei Xu, Xingkui Zhu 等NeurIPS 2024 · 被引用 380 次
- UST-SSM: Unified Spatio-Temporal State Space Models for Point Cloud Video ModelingPeiming Li, Ziyi Wang, Yulin Yuan, Hong Liu 等ICCV 2025 · 被引用 3 次
- PSTNet: Point Spatio-Temporal Convolution on Point Cloud SequencesHehe Fan, Xin Yu, Yuhang Ding, Yi Yang 等ICLR 2021 · 被引用 148 次
- DAPointMamba: Domain Adaptive Point Mamba for Point Cloud CompletionYinghui Li, Qianyu Zhou, Di Shao, Hao Yang 等AAAI 2026 · 被引用 1 次
