Foundation Models Defining A New Era In Sensor-based Human Activity Recognition: A Survey And Outlook
Sizhen Bian, Mengxi Liu, Lala Shakti Swarup Ray, Bo Zhou, Bin Guo, Zhiwen Yu, Thomas Plötz, Paul Lukowicz, Siyu Yuan, Vítor Fortes Rey
摘要
Sensor-based Human Activity Recognition (HAR) underpins many ubiquitous and wearable computing applications, yet current models remain limited by scarce labels, sensor heterogeneity, and weak generalization across users, devices, and contexts. Foundation models, which are generally pretrained at scale using self-supervised and multimodal learning, offer a unifying paradigm to address these challenges by learning reusable, adaptable representations for activity understanding. This survey synthesizes emerging foundation models for sensor-based HAR. We first clarify foundational concepts, definitions, and evaluation criteria, then organize existing work using a lifecycle-oriented taxonomy spanning input design, pretraining, adaptation, and utilization. Rather than enumerating individual models, we analyze recurring design patterns and trade-offs across nine technical axes, including modality scope, tokenization, architectures, learning paradigms, adaptation mechanisms, and deployment settings. From this synthesis, we identify three dominant development trajectories: (i) HAR-specific foundation models trained from scratch on large sensor corpora, (ii) adaptation of general time-series or multimodal foundation models to sensor-based HAR, and (iii) integration of large language models for reasoning, annotation, and human-AI interaction. We conclude by highlighting open challenges in data curation, multimodal alignment, personalization, privacy, and responsible deployment, and outline directions toward general-purpose, interpretable, and human-centered foundation models for activity understanding. A complete, continuously updated index of papers and models is available in our companion repository 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper41
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch 等ICML 2023 · 被引用 2,601 次
- A Time Series is Worth 64 Words: Long-term Forecasting with TransformersYuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, Jayant KalagnanamICLR 2023 · 被引用 536 次
相关 Paper
- Towards Customizable Foundation Models for Human Activity Recognition with Wearable DevicesMinghui Qiu, Cekai Weng, Mingming Fan, Kaishun WuUbiComp 2025 · 被引用 3 次
- Self-supervised Learning for Accelerometer-based Human Activity Recognition: A SurveyAleksej LogacjovUbiComp 2025 · 被引用 24 次
- MASTER: A Multi-modal Foundation Model for Human Activity RecognitionGuanzhou Zhu, Dong Zhao, Chunliang Li, Mingyue Zhao 等UbiComp 2025 · 被引用 8 次
- Wonderwall: A Virtual-to-Real Foundation Model for IMU-based HARShenghuan Miao, Ling ChenUbiComp 2026 · 被引用 2 次
- SensorLLM: Aligning Large Language Models with Motion Sensors for Human Activity RecognitionZechen Li, Shohreh Deldari, Linyao Chen, Hao Xue 等EMNLP 2025 · 被引用 9 次
