Multimodal Fusion via Teacher-Student Network for Indoor Action Recognition
Bruce X. B. Yu, Yan Liu, Keith C. C. Chan
Abstract
Indoor action recognition plays an important role in modern society, such as intelligent healthcare in large mobile cabin hospitals. With the wide usage of depth sensors like Kinect, multimodal information including skeleton and RGB modalities brings a promising way to improve the performance. However, existing methods are either focusing on a single data modality or failed to take the advantage of multiple data modalities. In this paper, we propose a Teacher-Student Multimodal Fusion (TSMF) model 1 that fuses the skeleton and RGB modalities at the model level for indoor action recognition. In our TSMF, we utilize a teacher network to transfer the structural knowledge of the skeleton modality to a student network for the RGB modality. With extensive experiments on two benchmarking datasets: NTU RGB+D and PKU-MMD, results show that the proposed TSMF consistently performs better than state-of-the-art single modal and multimodal methods. It also indicates that our TSMF could not only improve the accuracy of the student network but also significantly improve the ensemble accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7cc1c700-321e-41bf-8c1b-2cf1694262c4Cited by top-tier papers7
- GLA-GCN: Global-local Adaptive Graph Convolutional Network for 3D Human Pose Estimation from Monocular VideoBruce X. B. Yu, Zhi Zhang, Yongxu Liu, Sheng-Hua Zhong et al.ICCV 2023 · 131 citations
- Towards Good Practices for Missing Modality Robust Action RecognitionSangmin Woo, Sumin Lee, Yeonju Park, Muhammad Adi Nugroho et al.AAAI 2023 · 80 citations
- Cross-Modal Learning with 3D Deformable Attention for Action RecognitionSangwon Kim, Dasom Ahn, ByoungChul KoICCV 2023 · 49 citations
- Multi-Modality Co-Learning for Efficient Skeleton-based Action RecognitionJinfu Liu, Chen Chen, Mengyuan LiuACM MM 2024 · 27 citations
- Local-Global Multi-Modal Distillation for Weakly-Supervised Temporal Video GroundingPeijun Bao, Yong Xia, Wenhan Yang, Boon Poh Ng et al.AAAI 2024 · 20 citations
Builds on3
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Action Genome: Actions As Compositions of Spatio-Temporal Scene GraphsJingwei Ji, Ranjay Krishna, Li Fei-Fei, Juan Carlos NieblesCVPR 2020
- Disentangling and Unifying Graph Convolutions for Skeleton-Based Action RecognitionZiyu Liu, Hongwen Zhang, Zhenghao Chen, Zhiyong Wang et al.CVPR 2020
Related papers
- Lite-MKD: A Multi-modal Knowledge Distillation Framework for Lightweight Few-shot Action RecognitionBaolong Liu, Tianyi Zheng, Peng Zheng, Daizong Liu et al.ACM MM 2023 · 13 citations
- Skeletal Spatial-Temporal Semantics Guided Homogeneous-Heterogeneous Multimodal Network for Action RecognitionChenwei Zhang, Yuxuan Hu, Min Yang, Chengming Li et al.ACM MM 2023 · 4 citations
- HAAN: Human Action Aware Network for Multi-label Temporal Action DetectionZikai Gao, Peng Qiao, Yong DouACM MM 2023 · 7 citations
- Unified Multi-modal Unsupervised Representation Learning for Skeleton-based Action UnderstandingShengkai Sun, Daizong Liu, Jianfeng Dong, Xiaoye Qu et al.ACM MM 2023 · 34 citations
- Skeleton-based Action Recognition via Adaptive Cross-Form LearningXuanhan Wang, Yan Dai, Lianli Gao, Jingkuan SongACM MM 2022 · 28 citations
