IMUZero: Zero-Shot Human Activity Recognition by Language-Based Cross Modality Fusion
Jie Su, Fengtong Ge, Zhenyu Wen, Taotao Li, Yang Bai, Yejian Zhou, Xiaoqin Zhang
Abstract
Wearable-based human activity recognition (HAR) typically uses motion sensor data, such as inertial measurement unit (IMU) signals, to identify human movements. While effective in controlled scenarios, traditional HAR models are trained on a fixed set of activities and fail to generalize to new or unseen actions. This limitation motivates the use of zero-shot learning (ZSL), which aims to recognize unseen activities without direct training examples. Existing ZSL methods often rely on projecting seen and unseen classes into a shared latent space using external semantic information, such as visual or textual data. However, visual data are commonly unavailable in wearable settings, and text-based semantics from activity labels or coarse descriptions lack the detail needed for accurate recognition. Recent work explores large language models (LLMs) to provide prior knowledge through question-answering mechanisms. While promising, these approaches do not use raw sensor data directly and often miss important contextual signals. We propose IMUZero , a ZSL framework that fuses sensor signals with LLM-generated semantic attributes. Our method uses LLMs to produce fine-grained, decomposable activity attributes without additional LLM-based training, preserving sensor context. We also introduce a channel shuffle order constraint that models axial bias to improve generalization. Experiments on four public datasets show that our method outperforms existing ZSL approaches that rely on learned semantic embeddings. We release the code at https://github.com/Was-Lab/IMUZero.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ef71aa46-e123-45d3-b1b9-08999c1dd5efBuilds on12
- HSVA: Hierarchical Semantic-Visual Adaptation for Zero-Shot LearningShiming Chen, Guo-Sen Xie, Yang Liu, Qinmu Peng et al.NeurIPS 2021 · 190 citations
- TransZero: Attribute-Guided Transformer for Zero-Shot LearningShiming Chen, Ziming Hong, Yang Liu, Guo-Sen Xie et al.AAAI 2022 · 185 citations
- Compositional Zero-Shot Learning via Fine-Grained Dense Feature CompositionDat Huynh, Ehsan ElhamifarNeurIPS 2020 · 89 citations
- Making Sense of Sleep: Multimodal Sleep Stage Classification in a Large, Diverse Population Using Movement and Cardiac SensingBing Zhai, Ignacio Perez-Pozuelo, Emma A. D. Clifton, João R. M. Palotti et al.UbiComp 2020 · 79 citations
- IMUGPT 2.0: Language-Based Cross Modality Transfer for Sensor-Based Human Activity RecognitionZikang Leng, Amitrajit Bhattacharjee, Hrudhai Rajasekhar, Lizhe Zhang et al.UbiComp 2024 · 59 citations
Related papers
- ZeroHAR: Sensor Context Augments Zero-Shot Wearable Action RecognitionRanak Roy Chowdhury, Ritvik Kapila, Ameya Panse, Xiyuan Zhang et al.AAAI 2025 · 6 citations
- ZARA: Training-Free Motion Time-Series Reasoning via Evidence-Grounded LLM AgentsZechen Li, Baiyu Chen, Hao Xue, Flora D. SalimACL 2026
- Zero-Shot Learning for IMU-Based Activity Recognition Using Video EmbeddingsCatherine Tong, Jinchen Ge, Nicholas D. LaneUbiComp 2022 · 39 citations
- Large Language Model-guided Semantic Alignment for Human Activity RecognitionHua Yan, Heng Tan, Yi Ding, Pengfei Zhou et al.UbiComp 2026 · 3 citations
- SensorLLM: Aligning Large Language Models with Motion Sensors for Human Activity RecognitionZechen Li, Shohreh Deldari, Linyao Chen, Hao Xue et al.EMNLP 2025 · 9 citations
