Lune

UbiComp2025Top-tier venue

MASTER: A Multi-modal Foundation Model for Human Activity Recognition

Guanzhou Zhu, Dong Zhao, Chunliang Li, Mingyue Zhao, Zhengyuan Zhang, Hefeng Quan, Huadong Ma

2025Year
8Citations
2Top-tier citations

Abstract

Multi-modal sensing has become crucial in Human Activity Recognition (HAR) due to its ability to combine data from diverse sensors. However, challenges arise in recognizing various activities in different scenes using multi-modal data from different positions and devices, due to dynamic combinations of modal inputs, data heterogeneity, and scarcity of labeled data. To tackle these challenges, we propose MASTER, a multi-modal foundation model specifically designed for HAR. MASTER introduces a masked-data modeling-based self-supervised pre-training method, enabling the model to learn from unlabeled data and adapt to dynamic combinations of modal inputs. Moreover, it incorporates a few-shot alignment mechanism to facilitate adaptation to different activities, scenes, positions, and devices. Through the pre-training and fine-tuning on 7 multi-modal HAR datasets, MASTER currently supports, but is not limited to, 8 modalities (ACC, Gyro, mmWave, WiFi, Skeleton, Lidar, Infrared, and RGB) and 45 human activities. The results demonstrate that MASTER achieves the highest accuracy with minimal labeled data across various situations, surpassing alternative solutions.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get 35450740-e508-42d9-8976-c498d438151a

Cited by top-tier papers2

Ask how each one uses it

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines