Open-ended Human Activity Understanding via LLM-assisted Motion Decomposition and Semantic Fusion
Qingxin Wei, Kai Hu, Jiaming Huang, Cheng Guo, Yi Gao, Wei Dong
Abstract
Human activity is a fundamental component of intelligent interactive systems. Existing human activity recognition (HAR) follows a closed-ended assumption—defining a fixed set of activity classes and classifying each sensor signal into one of them—yet real-world human behavior is complex and ever-changing and thus cannot be fully represented by any fixed set. We present a new paradigm for open-ended human activity understanding (HAU), shifting the goal from closed-set classification to understanding activities via their meta-motions. To achieve this, we design MOSAIC, a motion decomposition and semantic fusion framework that transforms raw sensor streams into structured meta-motion representations and subsequently generates fine-grained natural-language descriptions. Additionally, we propose an LLM-driven training strategy for sensor-language alignment, enabling effective cross-modal supervision in the absence of real-world training data. Finally, we leverage LLM-assisted reasoning to recognize complex high-level activities based on the understanding of underlying motions and their composition. Our model demonstrates strong generalization and performance across 18 public HAR datasets, outperforming the best baseline by up to 19.6% in unseen scenarios.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- ZARA: Training-Free Motion Time-Series Reasoning via Evidence-Grounded LLM AgentsZechen Li, Baiyu Chen, Hao Xue, Flora D. SalimACL 2026
- SensorLLM: Aligning Large Language Models with Motion Sensors for Human Activity RecognitionZechen Li, Shohreh Deldari, Linyao Chen, Hao Xue et al.EMNLP 2025 · 9 citations
- IMUZero: Zero-Shot Human Activity Recognition by Language-Based Cross Modality FusionJie Su, Fengtong Ge, Zhenyu Wen, Taotao Li et al.UbiComp 2026 · 2 citations
- Large Language Model-guided Semantic Alignment for Human Activity RecognitionHua Yan, Heng Tan, Yi Ding, Pengfei Zhou et al.UbiComp 2026 · 3 citations
- MoPFormer: Motion-Primitive Transformer for Wearable-Sensor Activity RecognitionHao Zhang, Zhan Zhuang, Xuehao Wang, Xiaodong Yang et al.NeurIPS 2025 · 11 citations
