MM-Fit: Multimodal Deep Learning for Automatic Exercise Logging across Sensing Devices
David Strömbäck, Sangxia Huang, Valentin Radu
摘要
Fitness tracking devices have risen in popularity in recent years, but limitations in terms of their accuracy and failure to track many common exercises presents a need for improved fitness tracking solutions. This work proposes a multimodal deep learning approach to leverage multiple data sources for robust and accurate activity segmentation, exercise recognition and repetition counting. For this, we introduce the MM-Fit dataset; a substantial collection of inertial sensor data from smartphones, smartwatches and earbuds worn by participants while performing full-body workouts, and time-synchronised multi-viewpoint RGB-D video, with 2D and 3D pose estimates. We establish a strong baseline for activity segmentation and exercise recognition on the MM-Fit dataset, and demonstrate the effectiveness of our CNN-based architecture at extracting modality-specific spatial temporal features from inertial sensor and skeleton sequence data. We compare the performance of unimodal and multimodal models for activity recognition across a number of sensing devices and modalities. Furthermore, we demonstrate the effectiveness of multimodal deep learning at learning cross-modal representations for activity recognition, which achieves 96% accuracy across all sensing modalities on unseen subjects in the MM-Fit dataset; 94% using data from the smartwatch only; 85% from the smartphone only; and 82% on data from the earbud device. We strengthen single-device performance by using the zeroing-out training strategy, which phases out the other sensing modalities. Finally, we implement and evaluate a strong repetition counting baseline on our MM-Fit dataset. Collectively, these tasks contribute to recognising, segmenting and timing exercise and non-exercise activities for automatic exercise logging.
CCS Concepts: • Human-centered computing → Ubiquitous and mobile computing systems and tools; • Computing methodologies → Knowledge representation and reasoning; Neural networks; Learning latent representations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Sensing with Earables: A Systematic Literature Review and Taxonomy of PhenomenaTobias Röddiger, Christopher Clarke, Paula Breitling, Tim Schneegans 等UbiComp 2022 · 被引用 102 次
- WEAR: An Outdoor Sports Dataset for Wearable and Egocentric Activity RecognitionMarius Bock, Hilde Kuehne, Kristof Van Laerhoven, Michael MöllerUbiComp 2025 · 被引用 50 次
- ActSonic: Recognizing Everyday Activities from Inaudible Acoustic Wave Around the BodySaif Mahmud, Vineet Parikh, Qikang Liang, Ke Li 等UbiComp 2025 · 被引用 24 次
- OpenEarable 2.0: Open-Source Earphone Platform for Physiological Ear SensingTobias Röddiger, Michael Küttner, Philipp Lepold, Tobias King 等UbiComp 2025 · 被引用 19 次
- PhysiQ: Off-site Quality Assessment of Exercise in Physical TherapyHanchen David Wang, Meiyi MaUbiComp 2023 · 被引用 17 次
它引用的顶会 Paper1
相关 Paper
- SEGALL: A Unified Active Learning Framework for Wireless Sensing Data SegmentationNaiyu Zheng, Ruofeng Liu, Xiaoyi Fan, Cong Zhang 等UbiComp 2025 · 被引用 3 次
- SAMoSA: Sensing Activities with Motion and Subsampled AudioVimal Mollyn, Karan Ahuja, Dhruv Verma, Chris Harrison 等UbiComp 2022 · 被引用 54 次
- M3GYM: A Large-Scale Multimodal Multi-view Multi-person Pose Dataset for Fitness Activity Understanding in Real-world SettingsQingzheng Xu, Ru Cao, Xin Shen, Heming Du 等CVPR 2025
- MoBind: Motion Binding for Fine-Grained IMU-Video Pose AlignmentDuc Duy Nguyen, Tat-Jun Chin, Minh HoaiCVPR 2026 · 被引用 1 次
- IMUPoser: Full-Body Pose Estimation using IMUs in Phones, Watches, and EarbudsVimal Mollyn, Riku Arakawa, Mayank Goel, Chris Harrison 等CHI 2023 · 被引用 103 次
