Foundation Models Defining A New Era In Sensor-based Human Activity Recognition: A Survey And Outlook
Sizhen Bian, Mengxi Liu, Lala Shakti Swarup Ray, Bo Zhou, Bin Guo, Zhiwen Yu, Thomas Plötz, Paul Lukowicz, Siyu Yuan, Vítor Fortes Rey
Abstract
Sensor-based Human Activity Recognition (HAR) underpins many ubiquitous and wearable computing applications, yet current models remain limited by scarce labels, sensor heterogeneity, and weak generalization across users, devices, and contexts. Foundation models, which are generally pretrained at scale using self-supervised and multimodal learning, offer a unifying paradigm to address these challenges by learning reusable, adaptable representations for activity understanding. This survey synthesizes emerging foundation models for sensor-based HAR. We first clarify foundational concepts, definitions, and evaluation criteria, then organize existing work using a lifecycle-oriented taxonomy spanning input design, pretraining, adaptation, and utilization. Rather than enumerating individual models, we analyze recurring design patterns and trade-offs across nine technical axes, including modality scope, tokenization, architectures, learning paradigms, adaptation mechanisms, and deployment settings. From this synthesis, we identify three dominant development trajectories: (i) HAR-specific foundation models trained from scratch on large sensor corpora, (ii) adaptation of general time-series or multimodal foundation models to sensor-based HAR, and (iii) integration of large language models for reasoning, annotation, and human-AI interaction. We conclude by highlighting open challenges in data curation, multimodal alignment, personalization, privacy, and responsible deployment, and outline directions toward general-purpose, interpretable, and human-centered foundation models for activity understanding. A complete, continuously updated index of papers and models is available in our companion repository 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7c7de7cc-e954-402f-9ffa-6d20324daf37Builds on41
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch et al.ICML 2023 · 2,601 citations
- A Time Series is Worth 64 Words: Long-term Forecasting with TransformersYuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, Jayant KalagnanamICLR 2023 · 536 citations
Related papers
- Towards Customizable Foundation Models for Human Activity Recognition with Wearable DevicesMinghui Qiu, Cekai Weng, Mingming Fan, Kaishun WuUbiComp 2025 · 3 citations
- Self-supervised Learning for Accelerometer-based Human Activity Recognition: A SurveyAleksej LogacjovUbiComp 2025 · 24 citations
- MASTER: A Multi-modal Foundation Model for Human Activity RecognitionGuanzhou Zhu, Dong Zhao, Chunliang Li, Mingyue Zhao et al.UbiComp 2025 · 8 citations
- Wonderwall: A Virtual-to-Real Foundation Model for IMU-based HARShenghuan Miao, Ling ChenUbiComp 2026 · 2 citations
- SensorLLM: Aligning Large Language Models with Motion Sensors for Human Activity RecognitionZechen Li, Shohreh Deldari, Linyao Chen, Hao Xue et al.EMNLP 2025 · 9 citations
