GOAT: A Generalized Cross-Dataset Activity Recognition Framework with Natural Language Supervision
Shenghuan Miao, Ling Chen
Abstract
Wearable human activity recognition faces challenges in cross-dataset generalization due to variations in device configurations and activity types across datasets. We present GOAT, a Generalized crOss-dataset Activity recogniTion framework that leverages learning with natural language supervision to address these challenges. GOAT utilizes textual attributes from activity labels and device on-body positions to enable multimodal pre-training, aligning wearable activity representations with corresponding textual representations. This approach enables GOAT to adapt to diverse device configurations and activity label spaces in downstream tasks. Our method incorporates a novel device position encoding technique, a Transformer-based activity encoder, and a cosine similarity loss function to enhance feature extraction and generalization capabilities. Extensive evaluations demonstrate GOAT's effectiveness across various scenarios, including comparisons with state-of-the-art baselines, component analysis, and zero-shot activity recognition. GOAT shows promise for advancing cross-dataset activity recognition, offering a flexible and scalable solution for diverse wearable sensing applications.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 2680b3ed-99dc-4f01-94b9-0d982ca8b16bCited by top-tier papers3
- Large Language Model-guided Semantic Alignment for Human Activity RecognitionHua Yan, Heng Tan, Yi Ding, Pengfei Zhou et al.UbiComp 2026 · 3 citations
- Foundation Models Defining A New Era In Sensor-based Human Activity Recognition: A Survey And OutlookSizhen Bian, Mengxi Liu, Lala Shakti Swarup Ray, Bo Zhou et al.UbiComp 2026 · 2 citations
- DETACH : Decomposed Spatio-Temporal Alignment for Exocentric Video and Ambient Sensors with Staged LearningJunho Yoon, Jaemo Jeong, Hyunju Kim, Dongman LeeCVPR 2026
Related papers
- Limitations in Employing Natural Language Supervision for Sensor-Based Human Activity Recognition - And Ways to Overcome ThemHarish Haresamudram, Apoorva Beedu, Mashfiqui Rabbi, Sankalita Saha et al.AAAI 2025 · 11 citations
- SensorLM: Learning the Language of Wearable SensorsYuwei Zhang, Kumar Ayush, Siyuan Qiao, A. Ali Heydari et al.NeurIPS 2025 · 75 citations
- Wonderwall: A Virtual-to-Real Foundation Model for IMU-based HARShenghuan Miao, Ling ChenUbiComp 2026 · 2 citations
- MoPFormer: Motion-Primitive Transformer for Wearable-Sensor Activity RecognitionHao Zhang, Zhan Zhuang, Xuehao Wang, Xiaodong Yang et al.NeurIPS 2025 · 11 citations
- ZeroHAR: Sensor Context Augments Zero-Shot Wearable Action RecognitionRanak Roy Chowdhury, Ritvik Kapila, Ameya Panse, Xiyuan Zhang et al.AAAI 2025 · 6 citations
