Adapting Pretrained Large Vision Models for Sensor-based Activity Recognition
Yize Cai, Rui Feng, Kunlin Cai, Yunhuai Liu, Baoshen Guo, Zhiqing Hong
Abstract
Understanding and recognizing human activities from low-cost wearable sensors has attracted increasing attention in recent years. To achieve this goal, numerous learning models and augmentation approaches have been developed. While effective in certain scenarios, their performance is often limited due to the distribution shift between the training and testing data collected from different users or different body placements. Moreover, collecting large-scale and diverse sensor training data is costly and labor-intensive, which further exacerbates the data scarcity problem in HAR. To mitigate these challenges, in this work, we propose VisionHAR, a novel HAR framework that borrows knowledge from other data-rich modalities, i.e., low-cost and internet-scale image data. We first “draw” continuous sensor data on a figure to preserve both their temporal and periodic patterns. Then, we design a parameter-efficient transfer learning method to utilize the generalization capability of Large Vision Models (LVMs) pretrained on large-scale image data. To enable real-time activity recognition on edge devices, we further design a novel distillation approach to learn a highly effective and efficient student model, which achieves comparable performance with significantly fewer parameters. We evaluate our model in two main settings, i.e., cross-domain HAR and complex HAR (more than 18 activity categories). Experimental results show that VisionHAR outperforms the best existing HAR models by at least 8.22% in average accuracy and 11.48% in F1-score for cross-domain HAR with 1,460 times fewer parameters , and 7.62% in average accuracy and 7.86% in F1-score for complex HAR. Our findings suggest that the knowledge learned from vision data is generalizable to continuous sensor data, which provides a new potential solution for the data shortage issue in the activity recognition community. We release our code at https://github.com/saiketa/VisionHAR for future studies in the ubiquitous computing community.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get df9eaed7-4f94-41bf-aaa6-31185c8ee956Related papers
- CrossHAR: Generalizing Cross-dataset Human Activity Recognition via Hierarchical Self-Supervised PretrainingZhiqing Hong, Zelong Li, Shuxin Zhong, Wenjun Lyu et al.UbiComp 2024 · 64 citations
- SF-Adapter: Computational-Efficient Source-Free Domain Adaptation for Human Activity RecognitionHua Kang, Qingyong Hu, Qian ZhangUbiComp 2024 · 11 citations
- COMODO: Cross-Modal Video-to-IMU Distillation for Efficient Egocentric Human Activity RecognitionBaiyu Chen, Wilson Wongso, Zechen Li, Yonchanok Khaokaew et al.UbiComp 2026 · 1 citation
- Generalizable Low-Resource Activity Recognition with Diverse and Discriminative Representation LearningXin Qin, Jindong Wang, Shuo Ma, Wang Lu et al.KDD 2023 · 20 citations
- SelfHAR: Improving Human Activity Recognition through Self-training with Unlabeled DataChi Ian Tang, Ignacio Perez-Pozuelo, Dimitris Spathis, Søren Brage et al.UbiComp 2021 · 130 citations
