Towards Customizable Foundation Models for Human Activity Recognition with Wearable Devices
Minghui Qiu, Cekai Weng, Mingming Fan, Kaishun Wu
Abstract
Foundation models have achieved remarkable success across various domains by learning general representations from raw data, offering a promising paradigm for diverse applications. This concept holds great potential for advancing human activity recognition (HAR), particularly in overcoming challenges associated with collecting large-scale labeled datasets. However, the dynamic nature of HAR tasks, characterized by diverse sensing devices and activity types, results in fragmented datasets that question the feasibility of applying foundation model to this domain. In this work, we propose a novel foundation model training framework that effectively leverages heterogeneous datasets through a two-stage training strategy: (1) self-supervised learning to extract cross-domain sensor patterns, followed by (2) multi-task learning to align representations with semantic contexts. The effectiveness of the trained foundation model is demonstrated through extensive downstream experiments, with the superior fine-tuning performance across various modalities and input configurations-achieving the highest performance metric in 10 out of 12 settings-further validating the robustness and adaptability. While our model shows a performance gap compared to foundation models pre-trained on large-scale or high-quality data in zero-and few-shot scenarios, its competitive results with a more flexible architecture demonstrate the efficiency and potential of our training strategy for HAR foundation models.
CCS Concepts: • Human-centered computing → Ubiquitous and mobile computing systems and tools.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e704b410-1446-4099-8c8e-e45dae5a9978Cited by top-tier papers1
Ask how each one uses itBuilds on19
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu et al.NeurIPS 2020 · 1,957 citations
- One Fits All: Power General Time Series Analysis by Pretrained LMTian Zhou, Peisong Niu, Xue Wang, Liang Sun et al.NeurIPS 2023 · 1,178 citations
- Time-LLM: Time Series Forecasting by Reprogramming Large Language ModelsMing Jin, Shiyu Wang, Lintao Ma, Zhixuan Chu et al.ICLR 2024 · 915 citations
- TEMPO: Prompt-based Generative Pre-trained Transformer for Time Series ForecastingDefu Cao, Furong Jia, Sercan Ö. Arik, Tomas Pfister et al.ICLR 2024 · 262 citations
Related papers
- MASTER: A Multi-modal Foundation Model for Human Activity RecognitionGuanzhou Zhu, Dong Zhao, Chunliang Li, Mingyue Zhao et al.UbiComp 2025 · 8 citations
- Wonderwall: A Virtual-to-Real Foundation Model for IMU-based HARShenghuan Miao, Ling ChenUbiComp 2026 · 2 citations
- CrossHAR: Generalizing Cross-dataset Human Activity Recognition via Hierarchical Self-Supervised PretrainingZhiqing Hong, Zelong Li, Shuxin Zhong, Wenjun Lyu et al.UbiComp 2024 · 64 citations
- MobHAR: Source-free Knowledge Transfer for Human Activity Recognition on Mobile DevicesMeng Xue, Yinan Zhu, Wentao Xie, Zhixian Wang et al.UbiComp 2025 · 7 citations
- SelfHAR: Improving Human Activity Recognition through Self-training with Unlabeled DataChi Ian Tang, Ignacio Perez-Pozuelo, Dimitris Spathis, Søren Brage et al.UbiComp 2021 · 130 citations
