MASTER: A Multi-modal Foundation Model for Human Activity Recognition
Guanzhou Zhu, Dong Zhao, Chunliang Li, Mingyue Zhao, Zhengyuan Zhang, Hefeng Quan, Huadong Ma
Abstract
Multi-modal sensing has become crucial in Human Activity Recognition (HAR) due to its ability to combine data from diverse sensors. However, challenges arise in recognizing various activities in different scenes using multi-modal data from different positions and devices, due to dynamic combinations of modal inputs, data heterogeneity, and scarcity of labeled data. To tackle these challenges, we propose MASTER, a multi-modal foundation model specifically designed for HAR. MASTER introduces a masked-data modeling-based self-supervised pre-training method, enabling the model to learn from unlabeled data and adapt to dynamic combinations of modal inputs. Moreover, it incorporates a few-shot alignment mechanism to facilitate adaptation to different activities, scenes, positions, and devices. Through the pre-training and fine-tuning on 7 multi-modal HAR datasets, MASTER currently supports, but is not limited to, 8 modalities (ACC, Gyro, mmWave, WiFi, Skeleton, Lidar, Infrared, and RGB) and 45 human activities. The results demonstrate that MASTER achieves the highest accuracy with minimal labeled data across various situations, surpassing alternative solutions.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 35450740-e508-42d9-8976-c498d438151aCited by top-tier papers2
- Foundation Models Defining A New Era In Sensor-based Human Activity Recognition: A Survey And OutlookSizhen Bian, Mengxi Liu, Lala Shakti Swarup Ray, Bo Zhou et al.UbiComp 2026 · 2 citations
- SIGMA-ASL: Sensor-Integrated Multimodal Dataset for Sign Language RecognitionXiaofang Xiao, Guangchao Li, Guangrong Zhao, Qi Lin et al.UbiComp 2026
Related papers
- Towards Customizable Foundation Models for Human Activity Recognition with Wearable DevicesMinghui Qiu, Cekai Weng, Mingming Fan, Kaishun WuUbiComp 2025 · 3 citations
- CrossHAR: Generalizing Cross-dataset Human Activity Recognition via Hierarchical Self-Supervised PretrainingZhiqing Hong, Zelong Li, Shuxin Zhong, Wenjun Lyu et al.UbiComp 2024 · 64 citations
- X-Fi: A Modality-Invariant Foundation Model for Multimodal Human SensingXinyan Chen, Jianfei YangICLR 2025 · 1 citation
- SelfHAR: Improving Human Activity Recognition through Self-training with Unlabeled DataChi Ian Tang, Ignacio Perez-Pozuelo, Dimitris Spathis, Søren Brage et al.UbiComp 2021 · 130 citations
- Self-supervised Learning for Accelerometer-based Human Activity Recognition: A SurveyAleksej LogacjovUbiComp 2025 · 24 citations
