MMAct: A Large-Scale Dataset for Cross Modal Human Action Understanding
Quan Kong, Ziming Wu, Ziwei Deng, Martin Klinkigt, Bin Tong, Tomokazu Murakami
Abstract
Unlike vision modalities, body-worn sensors or passive sensing can avoid the failure of action understanding in vision related challenges, e.g. occlusion and appearance variation. However, a standard large-scale dataset does not exist, in which different types of modalities across vision and sensors are integrated. To address the disadvantage of vision-based modalities and push towards multi/cross modal action understanding, this paper introduces a new large-scale dataset recorded from 20 distinct subjects with seven different types of modalities: RGB videos, keypoints, acceleration, gyroscope, orientation, Wi-Fi and pressure signal. The dataset consists of more than 36k video clips for 37 action classes covering a wide range of daily life activities such as desktop-related and check-in-based ones in four different distinct scenarios. On the basis of our dataset, we propose a novel multi modality distillation model with attention mechanism to realize an adaptive knowledge transfer from sensor-based modalities to vision-based modali- ties. The proposed model significantly improves performance of action recognition compared to models trained with only RGB information. The experimental results confirm the effectiveness of our model on cross-subject, -view, -scene and -session evaluation criteria. We believe that this new large-scale multimodal dataset will contribute the community of multimodal based action understanding.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers21
- LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic AlignmentBin Zhu, Bin Lin, Munan Ning, Yang Yan et al.ICLR 2024 · 403 citations
- Cycle-Contrast for Self-Supervised Video Representation LearningQuan Kong, Wenpeng Wei, Ziwei Deng, Tomoaki Yoshinaga et al.NeurIPS 2020 · 59 citations
- MuMu: Cooperative Multitask Learning-Based Guided Multimodal FusionMd Mofijul Islam, Tariq IqbalAAAI 2022 · 56 citations
- UniMTS: Unified Pre-training for Motion Time SeriesXiyuan Zhang, Diyan Teng, Ranak Roy Chowdhury, Shuheng Li et al.NeurIPS 2024 · 49 citations
- Progressive Cross-modal Knowledge Distillation for Human Action RecognitionJianyuan Ni, Anne H. H. Ngu, Yan YanACM MM 2022 · 33 citations
Related papers
- MGR-Dark: A Large Multimodal Video Dataset and RGB-IR Benchmark for Gesture Recognition in DarknessYuanyuan Shi, Yunan Li, Siyu Liang, Huizhou Chen et al.ACM MM 2024 · 2 citations
- SATPose: Improving Monocular 3D Pose Estimation with Spatial-aware Ground TactilityLishuang Zhan, Enting Ying, Jiabao Gan, Shihui Guo et al.ACM MM 2024 · 2 citations
- DarkAct: A RGB-Thermal Dataset and Fusion Framework for Multimodal Low-Light Action RecognitionYuanjun Tan, Aoran Xiao, Liqian Deng, Zhigang TuCVPR 2026 · 1 citation
- Learning an Augmented RGB Representation with Cross-Modal Knowledge Distillation for Action DetectionRui Dai, Srijan Das, François BrémondICCV 2021 · 50 citations
- Adapting Pretrained Large Vision Models for Sensor-based Activity RecognitionYize Cai, Rui Feng, Kunlin Cai, Yunhuai Liu et al.UbiComp 2026
