Vi2ACT: Video-enhanced Cross-modal Co-learning with Representation Conditional Discriminator for Few-shot Human Activity Recognition
Kang Xia, Wenzhong Li, Yimiao Shao, Sanglu Lu
Abstract
Human Activity Recognition (HAR) as an emerging research field has attracted widespread academic attention due to its wide range of practical applications in areas such as healthcare, environmental monitoring, and sports training. Given the high cost of annotating sensor data, many unsupervised and semi-supervised methods have been applied to HAR to alleviate the problem of limited data. In this paper, we propose a novel video-enhanced cross-modal collaborative learning method, Vi2ACT, to address the issue of few-shot HAR. We introduce a new data augmentation approach that utilizes a text-to-video generation model to generate class-related videos. Subsequently, a large quantity of video semantic representations are obtained through fine-tuning the video encoder for cross-modal co-learning. Furthermore, to effectively align video semantic representations and time series representations, we enhance HAR at the representation-level using conditional Generative Adversarial Nets (cGAN). We design a novel Representation Conditional Discriminator that is trained to assess samples as originating from video representations rather than those generated by the time series encoder as accurately as possible. We conduct extensive experiments on four commonly used HAR datasets. The experimental results demonstrate that our method outperforms other baseline models in all few-shot scenarios.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get e1286bf9-9a3c-45dc-8f73-e991418d6d74Cited by top-tier papers1
Ask how each one uses itRelated papers
- TS2ACT: Few-Shot Human Activity Sensing with Cross-Modal Co-LearningKang Xia, Wenzhong Li, Shiwei Gan, Sanglu LuUbiComp 2024 · 18 citations
- VicTR: Video-conditioned Text Representations for Activity RecognitionKumara Kahatapitiya, Anurag Arnab, Arsha Nagrani, Michael S. RyooCVPR 2024
- Generalizable Low-Resource Activity Recognition with Diverse and Discriminative Representation LearningXin Qin, Jindong Wang, Shuo Ma, Wang Lu et al.KDD 2023 · 20 citations
- HMGAN: A Hierarchical Multi-Modal Generative Adversarial Network Model for Wearable Human Activity RecognitionLing Chen, Rong Hu, Menghan Wu, Xin ZhouUbiComp 2023 · 23 citations
- CDFSL-V: Cross-Domain Few-Shot Learning for VideosSarinda Samarasinghe, Mamshad Nayeem Rizve, Navid Kardan, Mubarak ShahICCV 2023 · 17 citations
