HMotionGPT: Aligning Hand Motions and Natural Language for Activity Understanding with Smart Rings
Yang Gao, Dong She, Wolin Liang, Chiyue Wang, Yingjing Xiao, Xianrong Yao, Cong Liu, Zhichao Huang, Zhanpeng Jin
Abstract
Hand-object interactions are central to everyday activities, yet most intelligent assistants today remain blind to users' physical actions. Existing IMU-based recognition approaches focus on classifying predefined gestures, but they lack the semantic expressiveness required for contextual support in real-world scenarios such as office work and home routines. In this paper, we introduce a semantic tokenization pipeline that bridges continuous inertial signals and large language models (LLMs), enabling assistants to “read” hand movements as naturally as words. We first collected a multimodal dataset of dual-hand activities across office and home environments capturing long-horizon action chains that span multiple interrelated sub-tasks. Using self-supervised representation learning, we discretize IMU embeddings into action tokens that approximate a vocabulary of hand interactions. These tokens are then aligned with natural language through instruction-tuned LLMs, supporting tasks such as action captioning, intent inference, and contextual feedback. Evaluation shows that our tokenization improves semantic consistency with language distributions, and the LLM produces accurate, human-preferred descriptions of actions across diverse activities. We further demonstrate a proof-of-concept assistant prototype that generates contextual reminders. Our findings highlight the potential of transforming raw hand motions into a “language of actions,” paving the way for everyday intelligent assistants that are aware of users' physical interactions. The Project page, source code, and dataset are publicly available at https://scut-hai.github.io/HMotionGPT/.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get f365a6de-ba01-4c05-89d9-973077a4a2aeRelated papers
- Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human InteractionsLiang Xu, Chengqun Yang, Zili Lin, Fei Xu et al.ICCV 2025 · 2 citations
- Human-Object Interaction from Human-level InstructionsZhen Wu, Jiaman Li, Pei Xu, C. Karen LiuICCV 2025 · 3 citations
- HOIGPT: Learning Long-Sequence Hand-Object Interaction with Language ModelsMingzhen Huang, Fu-Jen Chu, Bugra Tekin, Kevin J. Liang et al.CVPR 2025
- MotionMaster: Generalizable Text-Driven Motion Generation and EditingNan Jiang, Yunhao Li, Lexi Pang, Zimo He et al.CVPR 2026
- IMUGPT 2.0: Language-Based Cross Modality Transfer for Sensor-Based Human Activity RecognitionZikang Leng, Amitrajit Bhattacharjee, Hrudhai Rajasekhar, Lizhe Zhang et al.UbiComp 2024 · 59 citations
