TS2ACT: Few-Shot Human Activity Sensing with Cross-Modal Co-Learning
Kang Xia, Wenzhong Li, Shiwei Gan, Sanglu Lu
Abstract
Human Activity Recognition (HAR) based on embedded sensor data has become a popular research topic in ubiquitous computing, which has a wide range of practical applications in various fields such as human-computer interaction, healthcare, and motion tracking. Due to the difficulties of annotating sensing data, unsupervised and semi-supervised HAR methods are extensively studied, but their performance gap to the fully-supervised methods is notable. In this paper, we proposed a novel cross-modal co-learning approach called TS2ACT to achieve few-shot HAR. It introduces a cross-modal dataset augmentation method that uses the semantic-rich label text to search for human activity images to form an augmented dataset consisting of partially-labeled time series and fully-labeled images. Then it adopts a pre-trained CLIP image encoder to jointly train with a time series encoder using contrastive learning, where the time series and images are brought closer in feature space if they belong to the same activity class. For inference, the feature extracted from the input time series is compared with the embedding of a pre-trained CLIP text encoder using prompt learning, and the best match is output as the HAR classification results. We conducted extensive experiments on four public datasets to evaluate the performance of the proposed method. The numerical results show that TS2ACT significantly outperforms the state-of-the-art HAR methods, and it achieves performance close to or better than the fully supervised methods even using as few as 1% labeled data for model training. The source codes of TS2ACT are publicly available on GitHub1.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 6f123c63-e105-400c-a36d-06449274be6aCited by top-tier papers4
- Past, Present, and Future of Sensor-based Human Activity Recognition Using Wearables: A Surveying Tutorial on a Still Challenging TaskHarish Haresamudram, Chi Ian Tang, Sungho Suh, Paul Lukowicz et al.UbiComp 2025 · 31 citations
- Limitations in Employing Natural Language Supervision for Sensor-Based Human Activity Recognition - And Ways to Overcome ThemHarish Haresamudram, Apoorva Beedu, Mashfiqui Rabbi, Sankalita Saha et al.AAAI 2025 · 11 citations
- SensorLLM: Aligning Large Language Models with Motion Sensors for Human Activity RecognitionZechen Li, Shohreh Deldari, Linyao Chen, Hao Xue et al.EMNLP 2025 · 9 citations
- Towards Customizable Foundation Models for Human Activity Recognition with Wearable DevicesMinghui Qiu, Cekai Weng, Mingming Fan, Kaishun WuUbiComp 2025 · 3 citations
Related papers
- Vi2ACT: Video-enhanced Cross-modal Co-learning with Representation Conditional Discriminator for Few-shot Human Activity RecognitionKang Xia, Wenzhong Li, Yimiao Shao, Sanglu LuACM MM 2024 · 1 citation
- SelfHAR: Improving Human Activity Recognition through Self-training with Unlabeled DataChi Ian Tang, Ignacio Perez-Pozuelo, Dimitris Spathis, Søren Brage et al.UbiComp 2021 · 130 citations
- Cosmo: contrastive fusion learning with small data for multimodal human activity recognitionXiaomin Ouyang, Xian Shuai, Jiayu Zhou, Ivy Wang Shi et al.MobiCom 2022 · 94 citations
- Crossmodal Few-shot 3D Point Cloud Semantic SegmentationZiyu Zhao, Zhenyao Wu, Xinyi Wu, Canyu Zhang et al.ACM MM 2022 · 20 citations
- CrossHAR: Generalizing Cross-dataset Human Activity Recognition via Hierarchical Self-Supervised PretrainingZhiqing Hong, Zelong Li, Shuxin Zhong, Wenjun Lyu et al.UbiComp 2024 · 64 citations
