MM-Fit: Multimodal Deep Learning for Automatic Exercise Logging across Sensing Devices
David Strömbäck, Sangxia Huang, Valentin Radu
Abstract
Fitness tracking devices have risen in popularity in recent years, but limitations in terms of their accuracy and failure to track many common exercises presents a need for improved fitness tracking solutions. This work proposes a multimodal deep learning approach to leverage multiple data sources for robust and accurate activity segmentation, exercise recognition and repetition counting. For this, we introduce the MM-Fit dataset; a substantial collection of inertial sensor data from smartphones, smartwatches and earbuds worn by participants while performing full-body workouts, and time-synchronised multi-viewpoint RGB-D video, with 2D and 3D pose estimates. We establish a strong baseline for activity segmentation and exercise recognition on the MM-Fit dataset, and demonstrate the effectiveness of our CNN-based architecture at extracting modality-specific spatial temporal features from inertial sensor and skeleton sequence data. We compare the performance of unimodal and multimodal models for activity recognition across a number of sensing devices and modalities. Furthermore, we demonstrate the effectiveness of multimodal deep learning at learning cross-modal representations for activity recognition, which achieves 96% accuracy across all sensing modalities on unseen subjects in the MM-Fit dataset; 94% using data from the smartwatch only; 85% from the smartphone only; and 82% on data from the earbud device. We strengthen single-device performance by using the zeroing-out training strategy, which phases out the other sensing modalities. Finally, we implement and evaluate a strong repetition counting baseline on our MM-Fit dataset. Collectively, these tasks contribute to recognising, segmenting and timing exercise and non-exercise activities for automatic exercise logging.
CCS Concepts: • Human-centered computing → Ubiquitous and mobile computing systems and tools; • Computing methodologies → Knowledge representation and reasoning; Neural networks; Learning latent representations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 743565d4-5d52-4a07-b1f4-f9f9ea4f1998Cited by top-tier papers7
- Sensing with Earables: A Systematic Literature Review and Taxonomy of PhenomenaTobias Röddiger, Christopher Clarke, Paula Breitling, Tim Schneegans et al.UbiComp 2022 · 102 citations
- WEAR: An Outdoor Sports Dataset for Wearable and Egocentric Activity RecognitionMarius Bock, Hilde Kuehne, Kristof Van Laerhoven, Michael MöllerUbiComp 2025 · 50 citations
- ActSonic: Recognizing Everyday Activities from Inaudible Acoustic Wave Around the BodySaif Mahmud, Vineet Parikh, Qikang Liang, Ke Li et al.UbiComp 2025 · 24 citations
- OpenEarable 2.0: Open-Source Earphone Platform for Physiological Ear SensingTobias Röddiger, Michael Küttner, Philipp Lepold, Tobias King et al.UbiComp 2025 · 19 citations
- PhysiQ: Off-site Quality Assessment of Exercise in Physical TherapyHanchen David Wang, Meiyi MaUbiComp 2023 · 17 citations
Builds on1
Related papers
- SEGALL: A Unified Active Learning Framework for Wireless Sensing Data SegmentationNaiyu Zheng, Ruofeng Liu, Xiaoyi Fan, Cong Zhang et al.UbiComp 2025 · 3 citations
- SAMoSA: Sensing Activities with Motion and Subsampled AudioVimal Mollyn, Karan Ahuja, Dhruv Verma, Chris Harrison et al.UbiComp 2022 · 54 citations
- M3GYM: A Large-Scale Multimodal Multi-view Multi-person Pose Dataset for Fitness Activity Understanding in Real-world SettingsQingzheng Xu, Ru Cao, Xin Shen, Heming Du et al.CVPR 2025
- MoBind: Motion Binding for Fine-Grained IMU-Video Pose AlignmentDuc Duy Nguyen, Tat-Jun Chin, Minh HoaiCVPR 2026 · 1 citation
- IMUPoser: Full-Body Pose Estimation using IMUs in Phones, Watches, and EarbudsVimal Mollyn, Riku Arakawa, Mayank Goel, Chris Harrison et al.CHI 2023 · 103 citations
