Towards Customizable Foundation Models for Human Activity Recognition with Wearable Devices
Minghui Qiu, Cekai Weng, Mingming Fan, Kaishun Wu
摘要
Foundation models have achieved remarkable success across various domains by learning general representations from raw data, offering a promising paradigm for diverse applications. This concept holds great potential for advancing human activity recognition (HAR), particularly in overcoming challenges associated with collecting large-scale labeled datasets. However, the dynamic nature of HAR tasks, characterized by diverse sensing devices and activity types, results in fragmented datasets that question the feasibility of applying foundation model to this domain. In this work, we propose a novel foundation model training framework that effectively leverages heterogeneous datasets through a two-stage training strategy: (1) self-supervised learning to extract cross-domain sensor patterns, followed by (2) multi-task learning to align representations with semantic contexts. The effectiveness of the trained foundation model is demonstrated through extensive downstream experiments, with the superior fine-tuning performance across various modalities and input configurations-achieving the highest performance metric in 10 out of 12 settings-further validating the robustness and adaptability. While our model shows a performance gap compared to foundation models pre-trained on large-scale or high-quality data in zero-and few-shot scenarios, its competitive results with a more flexible architecture demonstrate the efficiency and potential of our training strategy for HAR foundation models.
CCS Concepts: • Human-centered computing → Ubiquitous and mobile computing systems and tools.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper19
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu 等NeurIPS 2020 · 被引用 1,957 次
- One Fits All: Power General Time Series Analysis by Pretrained LMTian Zhou, Peisong Niu, Xue Wang, Liang Sun 等NeurIPS 2023 · 被引用 1,178 次
- Time-LLM: Time Series Forecasting by Reprogramming Large Language ModelsMing Jin, Shiyu Wang, Lintao Ma, Zhixuan Chu 等ICLR 2024 · 被引用 915 次
- TEMPO: Prompt-based Generative Pre-trained Transformer for Time Series ForecastingDefu Cao, Furong Jia, Sercan Ö. Arik, Tomas Pfister 等ICLR 2024 · 被引用 262 次
相关 Paper
- MASTER: A Multi-modal Foundation Model for Human Activity RecognitionGuanzhou Zhu, Dong Zhao, Chunliang Li, Mingyue Zhao 等UbiComp 2025 · 被引用 8 次
- Wonderwall: A Virtual-to-Real Foundation Model for IMU-based HARShenghuan Miao, Ling ChenUbiComp 2026 · 被引用 2 次
- CrossHAR: Generalizing Cross-dataset Human Activity Recognition via Hierarchical Self-Supervised PretrainingZhiqing Hong, Zelong Li, Shuxin Zhong, Wenjun Lyu 等UbiComp 2024 · 被引用 64 次
- MobHAR: Source-free Knowledge Transfer for Human Activity Recognition on Mobile DevicesMeng Xue, Yinan Zhu, Wentao Xie, Zhixian Wang 等UbiComp 2025 · 被引用 7 次
- SelfHAR: Improving Human Activity Recognition through Self-training with Unlabeled DataChi Ian Tang, Ignacio Perez-Pozuelo, Dimitris Spathis, Søren Brage 等UbiComp 2021 · 被引用 130 次
