Decoupling and Recoupling Spatiotemporal Representation for RGB-D-based Motion Recognition
Benjia Zhou, Pichao Wang, Jun Wan, Yanyan Liang, Fan Wang, Du Zhang, Zhen Lei, Hao Li, Rong Jin
Abstract
Decoupling spatiotemporal representation refers to decomposing the spatial and temporal features into dimension-independent factors. Although previous RGB-D-based motion recognition methods have achieved promising performance through the tightly coupled multi-modal spatiotemporal representation, they still suffer from (i) optimization difficulty under small data setting due to the tightly spatiotemporal-entangled modeling; (ii) information redundancy as it usually contains lots of marginal information that is weakly relevant to classification; and (iii) low interaction between multi-modal spatiotemporal information caused by insufficient late fusion. To alleviate these drawbacks, we propose to decouple and recouple spatiotemporal representation for RGB-D-based motion recognition. Specifically, we disentangle the task of learning spatiotemporal representation into 3 sub-tasks: (1) Learning high-quality and dimension independent features through a decoupled spatial and temporal modeling network. (2) Recoupling the decoupled representation to establish stronger space-time dependency. (3) Introducing a Cross-modal Adaptive Posterior Fusion (CAPF) mechanism to capture cross-modal spatiotemporal information from RGB-D data. Seamless combination of these novel designs forms a robust spatiotemporal representation and achieves better performance than state-of-the-art methods on four public motion datasets. Our code is available at https://github.com/damo-cv/MotionRGBD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0dd1fcc1-0d99-48e0-b517-460412f4a4eaCited by top-tier papers6
- Gloss-free Sign Language Translation: Improving from Visual-Language PretrainingBenjia Zhou, Zhigang Chen, Albert Clapés, Jun Wan et al.ICCV 2023 · 123 citations
- Mitigating and Evaluating Static Bias of Action Representations in the Background and the ForegroundHaoxin Li, Yuan Liu, Hanwang Zhang, Boyang LiICCV 2023 · 30 citations
- Multi-stage Factorized Spatio-Temporal Representation for RGB-D Action and Gesture RecognitionYujun Ma, Benjia Zhou, Ruili Wang, Pichao WangACM MM 2023 · 21 citations
- UMIFormer: Mining the Correlations between Similar Tokens for Multi-View 3D ReconstructionZhenwei Zhu, Liying Yang, Ning Li, Chaohao Jiang et al.ICCV 2023 · 12 citations
- Learning Robust Representations with Information Bottleneck and Memory Network for RGB-D-based Gesture RecognitionYunan Li, Huizhou Chen, Guanwen Feng, Qiguang MiaoICCV 2023 · 10 citations
Builds on8
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Channel-wise Topology Refinement Graph Convolution for Skeleton-Based Action RecognitionYuxin Chen, Ziqi Zhang, Chunfeng Yuan, Bing Li et al.ICCV 2021 · 871 citations
- Regional Attention with Architecture-Rebuilt 3D Network for RGB-D Gesture RecognitionBenjia Zhou, Yunan Li, Jun WanAAAI 2021 · 32 citations
- An Efficient PointLSTM for Point Clouds Based Gesture RecognitionYuecong Min, Yanxiao Zhang, Xiujuan Chai, Xilin ChenCVPR 2020
Related papers
- Attribute-Based Progressive Fusion Network for RGBT TrackingYun Xiao, Mengmeng Yang, Chenglong Li, Lei Liu et al.AAAI 2022 · 218 citations
- Beyond Duality: A Hybrid Framework of Leveraging Shared and Private Features for RGB-Event Object DetectionKeyao Wang, Shuai Liu, Hengda Shi, Lukui Shi et al.CVPR 2026
- Not All Frame Features are Equal: Video-to-4D Generation via Decoupling Dynamic-Static FeaturesLiying Yang, Chen Liu, Zhenwei Zhu, Ajian Liu et al.ICCV 2025
- End-to-End RGB-D Image Compression via Exploiting Channel-Modality RedundancyHuiming Zheng, Wei GaoAAAI 2024 · 15 citations
- Appearance-Motion Decomposed Alignment for Text-Video RetrievalMeng Meng, Zichang Tan, Yong Zhang, Xu ZhouAAAI 2026
