ZeroHAR: Sensor Context Augments Zero-Shot Wearable Action Recognition
Ranak Roy Chowdhury, Ritvik Kapila, Ameya Panse, Xiyuan Zhang, Diyan Teng, Rashmi Kulkarni, Dezhi Hong, Rajesh K. Gupta, Jingbo Shang
Abstract
Wearable Human Action Recognition (wHAR) uses motion sensor data to identify human movements, which is essential for mobile and wearable devices. However, traditional wHAR systems are only trained on a limited set of activities. Hence, they fail to generalize to diverse human motions, prompting Zero-Shot Learning (ZSL). Existing ZSL methods for wHAR focus solely on augmenting labels, such as representing them as attribute matrices, images, videos, or text. We propose ZeroHAR that enhances ZSL by not just focusing on activity labels, but also augmenting motion data with sensor context features. Our approach incorporates information about the sensor type, the Cartesian axis of the data, and the sensor's body position, providing the model with crucial spatial and biomechanical insights. This helps the model to generalize better to new actions. First, we train the model by aligning the latent space of the motion time series with its corresponding sensor context, while distancing it from unrelated sensor contexts. Then, we train the model using the target activity descriptions. We tested our method against eight baselines on five benchmark HAR datasets with various sensors, placements, and activities. Our model shows exceptional generalizability across the 18 motion time series classification benchmark datasets, outperforming the best baselines by 262% in the zero-shot setting.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- SensorChat: Answering Qualitative and Quantitative Questions during Long-term Multimodal Sensor InteractionsXiaofan Yu, Lanxiang Hu, Benjamin Z. Reichman, Dylan Chu et al.UbiComp 2025 · 4 citations
- ZARA: Training-Free Motion Time-Series Reasoning via Evidence-Grounded LLM AgentsZechen Li, Baiyu Chen, Hao Xue, Flora D. SalimACL 2026
Builds on6
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- MMAct: A Large-Scale Dataset for Cross Modal Human Action UnderstandingQuan Kong, Ziming Wu, Ziwei Deng, Martin Klinkigt et al.ICCV 2019 · 108 citations
- A Transformer-based Framework for Multivariate Time Series Representation LearningGeorge Zerveas, Srideepika Jayaraman, Dhaval Patel, Anuradha Bhamidipaty et al.KDD 2021 · 66 citations
Related papers
- IMUZero: Zero-Shot Human Activity Recognition by Language-Based Cross Modality FusionJie Su, Fengtong Ge, Zhenyu Wen, Taotao Li et al.UbiComp 2026 · 2 citations
- UniMTS: Unified Pre-training for Motion Time SeriesXiyuan Zhang, Diyan Teng, Ranak Roy Chowdhury, Shuheng Li et al.NeurIPS 2024 · 49 citations
- Zero-Shot Learning for IMU-Based Activity Recognition Using Video EmbeddingsCatherine Tong, Jinchen Ge, Nicholas D. LaneUbiComp 2022 · 39 citations
- Augmented Adversarial Learning for Human Activity Recognition with Partial Sensor SetsHua Kang, Qianyi Huang, Qian ZhangUbiComp 2022 · 16 citations
- Wonderwall: A Virtual-to-Real Foundation Model for IMU-based HARShenghuan Miao, Ling ChenUbiComp 2026 · 2 citations
