ZeroHAR: Sensor Context Augments Zero-Shot Wearable Action Recognition
Ranak Roy Chowdhury, Ritvik Kapila, Ameya Panse, Xiyuan Zhang, Diyan Teng, Rashmi Kulkarni, Dezhi Hong, Rajesh K. Gupta, Jingbo Shang
摘要
Wearable Human Action Recognition (wHAR) uses motion sensor data to identify human movements, which is essential for mobile and wearable devices. However, traditional wHAR systems are only trained on a limited set of activities. Hence, they fail to generalize to diverse human motions, prompting Zero-Shot Learning (ZSL). Existing ZSL methods for wHAR focus solely on augmenting labels, such as representing them as attribute matrices, images, videos, or text. We propose ZeroHAR that enhances ZSL by not just focusing on activity labels, but also augmenting motion data with sensor context features. Our approach incorporates information about the sensor type, the Cartesian axis of the data, and the sensor's body position, providing the model with crucial spatial and biomechanical insights. This helps the model to generalize better to new actions. First, we train the model by aligning the latent space of the motion time series with its corresponding sensor context, while distancing it from unrelated sensor contexts. Then, we train the model using the target activity descriptions. We tested our method against eight baselines on five benchmark HAR datasets with various sensors, placements, and activities. Our model shows exceptional generalizability across the 18 motion time series classification benchmark datasets, outperforming the best baselines by 262% in the zero-shot setting.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- SensorChat: Answering Qualitative and Quantitative Questions during Long-term Multimodal Sensor InteractionsXiaofan Yu, Lanxiang Hu, Benjamin Z. Reichman, Dylan Chu 等UbiComp 2025 · 被引用 4 次
- ZARA: Training-Free Motion Time-Series Reasoning via Evidence-Grounded LLM AgentsZechen Li, Baiyu Chen, Hao Xue, Flora D. SalimACL 2026
它引用的顶会 Paper6
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- MMAct: A Large-Scale Dataset for Cross Modal Human Action UnderstandingQuan Kong, Ziming Wu, Ziwei Deng, Martin Klinkigt 等ICCV 2019 · 被引用 108 次
- A Transformer-based Framework for Multivariate Time Series Representation LearningGeorge Zerveas, Srideepika Jayaraman, Dhaval Patel, Anuradha Bhamidipaty 等KDD 2021 · 被引用 66 次
相关 Paper
- IMUZero: Zero-Shot Human Activity Recognition by Language-Based Cross Modality FusionJie Su, Fengtong Ge, Zhenyu Wen, Taotao Li 等UbiComp 2026 · 被引用 2 次
- UniMTS: Unified Pre-training for Motion Time SeriesXiyuan Zhang, Diyan Teng, Ranak Roy Chowdhury, Shuheng Li 等NeurIPS 2024 · 被引用 49 次
- Zero-Shot Learning for IMU-Based Activity Recognition Using Video EmbeddingsCatherine Tong, Jinchen Ge, Nicholas D. LaneUbiComp 2022 · 被引用 39 次
- Augmented Adversarial Learning for Human Activity Recognition with Partial Sensor SetsHua Kang, Qianyi Huang, Qian ZhangUbiComp 2022 · 被引用 16 次
- Wonderwall: A Virtual-to-Real Foundation Model for IMU-based HARShenghuan Miao, Ling ChenUbiComp 2026 · 被引用 2 次
