MultiHGR: Multi-Task Hand Gesture Recognition with Cross-Modal Wrist-Worn Devices
Mengxia Lyu, Hao Zhou, Kaiwen Guo, Wangqiu Zhou, Xingfa Shen, Yu Gu
摘要
Hand gesture recognition (HGR) is essential for human-machine interaction. Although the existing solutions achieve good performance in specific tasks, they still face challenges when users navigate through different application contexts, i.e., demanding multi-task ability to support newly arrived HGR tasks. In this paper, we propose the first IMU-vision based system hosted on wrist-worn devices to support multi-task HGR, denoted as MultiHGR. The system introduces a novel two-stage training strategy, i.e., task-agnostic stage to align cross-modal features from unlabeled arbitrary gesture through contrastive learning, and task-related stage to learn modality contributions with limited labeled data in specific tasks through self-attention mechanism. Since only the second task-related stage should be executed for each new task, MultiHGR could accommodate multiple tasks with significant reduced training cost and storage requirement. The evaluation results on three HGR tasks demonstrates that MultiHGR reduces 64.92% training time, and 24.04% storage as compared with traditional multimodal single-task models, and MultiHGR outperforms unimodal single-task models with 14.37%, 19.28%, and 31% improvements in these three tasks, respectively. As compared with state-of-the-art multimodal single-task model, MultiHGR achieves average 6.35% accuracy improvement, along with 65.74% training time reduction.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Mudra: A Multi-Modal Smartwatch Interactive System with Hand Gesture Recognition and User IdentificationKaiwen Guo, Hao Zhou, Ye Tian, Wangqiu Zhou 等INFOCOM 2022 · 被引用 18 次
- iRadar: Synthesizing Millimeter-Waves from Wearable Inertial Inputs for Human Gesture SensingHuanqi Yang, Mingda Han, Xinyue Li, Di Duan 等INFOCOM 2025 · 被引用 9 次
- BiCrossNet with Decoupled Dual Generators: A Parameter‑Efficient and Generalizable Few‑Shot Custom Gesture Recognition Frameworkchunyang yu, Ning SUN, Shenyue WangICML 2026
- Synthetic Smartwatch IMU Data Generation from In-the-wild ASL VideosPanneer Selvam Santhalingam, Parth Pathak, Huzefa Rangwala, Jana KoseckaUbiComp 2023 · 被引用 28 次
- Cosmo: contrastive fusion learning with small data for multimodal human activity recognitionXiaomin Ouyang, Xian Shuai, Jiayu Zhou, Ivy Wang Shi 等MobiCom 2022 · 被引用 94 次
