Zero-Shot Learning for IMU-Based Activity Recognition Using Video Embeddings
Catherine Tong, Jinchen Ge, Nicholas D. Lane
Abstract
The Activity Recognition Chain generally precludes the challenging scenario of recognizing new activities that were unseen during training, despite this scenario being a practical and common one as users perform diverse activities at test time. A few prior works have adopted zero-shot learning methods for IMU-based activity recognition, which work by relating seen and unseen classes through an auxiliary semantic space. However, these methods usually rely heavily on a hand-crafted attribute space which is costly to define, or a learnt semantic space based on word embedding, which lacks motion-related information crucial for distinguishing IMU features. Instead, we propose a strategy to exploit videos of human activities to construct an informative semantic space. With our approach, knowledge from state-of-the-art video action recognition models is encoded into video embeddings to relate seen and unseen activity classes. Experiments on three public datasets find that our approach outperforms other learnt semantic spaces, with an additional desirable feature of scalability, as recognition performance is seen to scale with the amount of data used. More generally, our results indicate that exploiting information from the video domain for IMU-based tasks is a promising direction, with tangible returns in a zero-shot learning scenario.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4d4aac2e-b692-41b0-915b-b3cee64d4f2eCited by top-tier papers10
- EgoDistill: Egocentric Head Motion Distillation for Efficient Video UnderstandingShuhan Tan, Tushar Nagarajan, Kristen GraumanNeurIPS 2023 · 44 citations
- Synthetic Smartwatch IMU Data Generation from In-the-wild ASL VideosPanneer Selvam Santhalingam, Parth Pathak, Huzefa Rangwala, Jana KoseckaUbiComp 2023 · 28 citations
- MMTSA: Multi-Modal Temporal Segment Attention Network for Efficient Human Activity RecognitionZiqi Gao, Yuntao Wang, Jianguo Chen, Junliang Xing et al.UbiComp 2023 · 22 citations
- Taming Event Cameras with Bio-Inspired Architecture and Algorithm: A Case for Drone Obstacle AvoidanceJingao Xu, Danyang Li, Zheng Yang, Yishujie Zhao et al.MobiCom 2023 · 14 citations
- Limitations in Employing Natural Language Supervision for Sensor-Based Human Activity Recognition - And Ways to Overcome ThemHarish Haresamudram, Apoorva Beedu, Mashfiqui Rabbi, Sankalita Saha et al.AAAI 2025 · 11 citations
Builds on4
- IMUTube: Automatic Extraction of Virtual on-body Accelerometry from Video for Human Activity RecognitionHyeokHyen Kwon, Catherine Tong, Harish Haresamudram, Yan Gao et al.UbiComp 2020 · 153 citations
- Vid2Doppler: Synthesizing Doppler Radar Data from Videos for Training Privacy-Preserving Activity RecognitionKaran Ahuja, Yue Jiang, Mayank Goel, Chris HarrisonCHI 2021 · 118 citations
- Approaching the Real-World: Supporting Activity Recognition Training with Virtual IMU DataHyeokHyen Kwon, Bingyao Wang, Gregory D. Abowd, Thomas PlötzUbiComp 2021 · 48 citations
- Teaching RF to Sense without RF Training MeasurementsHong Cai, Belal Korany, Chitra R. Karanam, Yasamin MostofiUbiComp 2021 · 42 citations
Related papers
- IMUZero: Zero-Shot Human Activity Recognition by Language-Based Cross Modality FusionJie Su, Fengtong Ge, Zhenyu Wen, Taotao Li et al.UbiComp 2026 · 2 citations
- Zero-shot Skeleton-based Action Recognition via Mutual Information Estimation and MaximizationYujie Zhou, Wenwen Qiang, Anyi Rao, Ning Lin et al.ACM MM 2023 · 25 citations
- ZeroHAR: Sensor Context Augments Zero-Shot Wearable Action RecognitionRanak Roy Chowdhury, Ritvik Kapila, Ameya Panse, Xiyuan Zhang et al.AAAI 2025 · 6 citations
- Disentangling Visual Embeddings for Attributes and ObjectsNirat Saini, Khoi Pham, Abhinav ShrivastavaCVPR 2022 · 74 citations
- VGSE: Visually-Grounded Semantic Embeddings for Zero-Shot LearningWenjia Xu, Yongqin Xian, Jiuniu Wang, Bernt Schiele et al.CVPR 2022 · 61 citations
