MMAct: A Large-Scale Dataset for Cross Modal Human Action Understanding
Quan Kong, Ziming Wu, Ziwei Deng, Martin Klinkigt, Bin Tong, Tomokazu Murakami
摘要
Unlike vision modalities, body-worn sensors or passive sensing can avoid the failure of action understanding in vision related challenges, e.g. occlusion and appearance variation. However, a standard large-scale dataset does not exist, in which different types of modalities across vision and sensors are integrated. To address the disadvantage of vision-based modalities and push towards multi/cross modal action understanding, this paper introduces a new large-scale dataset recorded from 20 distinct subjects with seven different types of modalities: RGB videos, keypoints, acceleration, gyroscope, orientation, Wi-Fi and pressure signal. The dataset consists of more than 36k video clips for 37 action classes covering a wide range of daily life activities such as desktop-related and check-in-based ones in four different distinct scenarios. On the basis of our dataset, we propose a novel multi modality distillation model with attention mechanism to realize an adaptive knowledge transfer from sensor-based modalities to vision-based modali- ties. The proposed model significantly improves performance of action recognition compared to models trained with only RGB information. The experimental results confirm the effectiveness of our model on cross-subject, -view, -scene and -session evaluation criteria. We believe that this new large-scale multimodal dataset will contribute the community of multimodal based action understanding.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic AlignmentBin Zhu, Bin Lin, Munan Ning, Yang Yan 等ICLR 2024 · 被引用 403 次
- Cycle-Contrast for Self-Supervised Video Representation LearningQuan Kong, Wenpeng Wei, Ziwei Deng, Tomoaki Yoshinaga 等NeurIPS 2020 · 被引用 59 次
- MuMu: Cooperative Multitask Learning-Based Guided Multimodal FusionMd Mofijul Islam, Tariq IqbalAAAI 2022 · 被引用 56 次
- UniMTS: Unified Pre-training for Motion Time SeriesXiyuan Zhang, Diyan Teng, Ranak Roy Chowdhury, Shuheng Li 等NeurIPS 2024 · 被引用 49 次
- Progressive Cross-modal Knowledge Distillation for Human Action RecognitionJianyuan Ni, Anne H. H. Ngu, Yan YanACM MM 2022 · 被引用 33 次
相关 Paper
- MGR-Dark: A Large Multimodal Video Dataset and RGB-IR Benchmark for Gesture Recognition in DarknessYuanyuan Shi, Yunan Li, Siyu Liang, Huizhou Chen 等ACM MM 2024 · 被引用 2 次
- SATPose: Improving Monocular 3D Pose Estimation with Spatial-aware Ground TactilityLishuang Zhan, Enting Ying, Jiabao Gan, Shihui Guo 等ACM MM 2024 · 被引用 2 次
- DarkAct: A RGB-Thermal Dataset and Fusion Framework for Multimodal Low-Light Action RecognitionYuanjun Tan, Aoran Xiao, Liqian Deng, Zhigang TuCVPR 2026 · 被引用 1 次
- Learning an Augmented RGB Representation with Cross-Modal Knowledge Distillation for Action DetectionRui Dai, Srijan Das, François BrémondICCV 2021 · 被引用 50 次
- Adapting Pretrained Large Vision Models for Sensor-based Activity RecognitionYize Cai, Rui Feng, Kunlin Cai, Yunhuai Liu 等UbiComp 2026
