Spatial-Temporal Masked Autoencoder for Multi-Device Wearable Human Activity Recognition
Shenghuan Miao, Ling Chen, Rong Hu
Abstract
The widespread adoption of wearable devices has led to a surge in the development of multi-device wearable human activity recognition (WHAR) systems. Nevertheless, the performance of traditional supervised learning-based methods to WHAR is limited by the challenge of collecting ample annotated wearable data. To overcome this limitation, self-supervised learning (SSL) has emerged as a promising solution by first training a competent feature extractor on a substantial quantity of unlabeled data, followed by refining a minimal classifier with a small amount of labeled data. Despite the promise of SSL in WHAR, the majority of studies have not considered missing device scenarios in multi-device WHAR. To bridge this gap, we propose a multi-device SSL WHAR method termed Spatial-Temporal Masked Autoencoder (STMAE). STMAE captures discriminative activity representations by utilizing the asymmetrical encoder-decoder structure and two-stage spatial-temporal masking strategy, which can exploit the spatial-temporal correlations in multi-device data to improve the performance of SSL WHAR, especially on missing device scenarios. Experiments on four real-world datasets demonstrate the efficacy of STMAE in various practical scenarios.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get cd54a1b7-354f-4e3b-835f-b0ba65aab4eeCited by top-tier papers5
- Past, Present, and Future of Sensor-based Human Activity Recognition Using Wearables: A Surveying Tutorial on a Still Challenging TaskHarish Haresamudram, Chi Ian Tang, Sungho Suh, Paul Lukowicz et al.UbiComp 2025 · 31 citations
- Foundation Models Defining A New Era In Sensor-based Human Activity Recognition: A Survey And OutlookSizhen Bian, Mengxi Liu, Lala Shakti Swarup Ray, Bo Zhou et al.UbiComp 2026 · 2 citations
- COMODO: Cross-Modal Video-to-IMU Distillation for Efficient Egocentric Human Activity RecognitionBaiyu Chen, Wilson Wongso, Zechen Li, Yonchanok Khaokaew et al.UbiComp 2026 · 1 citation
- MODA: Motion-Drift Augmentation for Inertial Human Motion AnalysisYinghao Wu, Shihui Guo, Yipeng QinCVPR 2025
- CURE: Context-driven Diffusion with Progressive Expansion for Single Domain Generalization in Time Series ClassificationYuhang Pei, Fanchun Meng, Wenrui Wu, Tao Ren et al.ICML 2026
Related papers
- Self-supervised Learning for Accelerometer-based Human Activity Recognition: A SurveyAleksej LogacjovUbiComp 2025 · 24 citations
- Assessing the State of Self-Supervised Human Activity Recognition Using WearablesHarish Haresamudram, Irfan Essa, Thomas PlötzUbiComp 2022 · 104 citations
- ColloSSL: Collaborative Self-Supervised Learning for Human Activity RecognitionYash Jain, Chi Ian Tang, Chulhong Min, Fahim Kawsar et al.UbiComp 2022 · 113 citations
- HiMAE: Hierarchical Masked Autoencoders Discover Resolution-Specific Structure in Wearable Time SeriesSimon A. Lee, Cyrus Tanade, Hao Zhou, Juhyeon Lee et al.ICLR 2026 · 22 citations
- Guiding Masked Representation Learning to Capture Spatio-Temporal Relationship of ElectrocardiogramYeongyeon Na, Minje Park, Yunwon Tae, Sunghoon JooICLR 2024 · 92 citations
