AIFit: Automatic 3D Human-Interpretable Feedback Models for Fitness Training
Mihai Fieraru, Mihai Zanfir, Silviu Cristian Pirlea, Vlad Olaru, Cristian Sminchisescu
Abstract
I went to the gym today, but how well did I do? And where should I improve? Ah, my back hurts slightly... User engagement can be sustained and injuries avoided by being able to reconstruct 3d human pose and motion, relate it to good training practices, identify errors, and provide early, real-time feedback. In this paper we introduce the first automatic system, AIFit, that performs 3d human sensing for fitness training. The system can be used at home, outdoors, or at the gym. AIFit is able to reconstruct 3d human pose, shape, and motion, reliably segment exercise repetitions, and identify in real-time the deviations between standards learnt from trainers, and the execution of a trainee. As a result, localized, quantitative feedback for correct execution of exercises, reduced risk of injury, and continuous improvement is possible. To support research and evaluation, we introduce the first large scale dataset, Fit3D, containing over 3 million images and corresponding 3d human shape and motion capture ground truth configurations, with over 37 repeated exercises, covering all the major muscle groups, performed by instructors and trainees. Our statistical coach is governed by a global parameter that captures how critical it should be of a trainee's performance. This is an important aspect that helps adapt to a student's level of fitness (i.e. beginner vs. advanced vs. expert), or to the expected accuracy of a 3d pose reconstruction method. We show that, for different values of the global parameter, our feedback system based on 3d pose estimates achieves good accuracy compared to the one based on ground-truth motion capture. Our statistical coach offers feedback in natural language, and with spatio-temporal visual grounding.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 605e6f91-8ccc-423a-a32b-d0db5e22541fCited by top-tier papers20
- Video Action DifferencingJames Burgess, Xiaohan Wang, Yuhui Zhang, Anita Rau et al.ICLR 2025 · 1,149 citations
- Neural Localizer Fields for Continuous 3D Human Pose and Shape EstimationIstván Sárándi, Gerard Pons-MollNeurIPS 2024 · 76 citations
- PoseFix: Correcting 3D Human Poses with Natural LanguageGinger Delmas, Philippe Weinzaepfel, Francesc Moreno-Noguer, Grégory RogezICCV 2023 · 49 citations
- H3WB: Human3.6M 3D WholeBody Dataset and BenchmarkYue Zhu, Nermin Samet, David PicardICCV 2023 · 34 citations
- The Quest for Generalizable Motion Generation: Data, Model, and EvaluationJing Lin, Ruisi Wang, Junzhe Lu, Ziqi Huang et al.ICLR 2026 · 23 citations
Builds on7
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 1,139 citations
- PIE: A Large-Scale Dataset and Models for Pedestrian Intention Estimation and Trajectory PredictionAmir Rasouli, Iuliia Kotseruba, Toni Kunic, John K. TsotsosICCV 2019 · 411 citations
- Mask-Guided Attention Network for Occluded Pedestrian DetectionYanwei Pang, Jin Xie, Muhammad Haris Khan, Rao Muhammad Anwer et al.ICCV 2019 · 216 citations
- Learning Complex 3D Human Self-ContactMihai Fieraru, Mihai Zanfir, Elisabeta Oneata, Alin-Ionut Popa et al.AAAI 2021 · 44 citations
- Counting Out Time: Class Agnostic Video Repetition Counting in the WildDebidatta Dwibedi, Yusuf Aytar, Jonathan Tompson, Pierre Sermanet et al.CVPR 2020
Related papers
- FLAG3D: A 3D Fitness Activity Dataset with Language InstructionYansong Tang, Jinpeng Liu, Aoyang Liu, Bin Yang et al.CVPR 2023
- M3GYM: A Large-Scale Multimodal Multi-view Multi-person Pose Dataset for Fitness Activity Understanding in Real-world SettingsQingzheng Xu, Ru Cao, Xin Shen, Heming Du et al.CVPR 2025
- Real-time Multi-modal Comprehensive Bodyweight Exercise Feedback System with a Personal Virtual Expert GuidingYunho Choi, Jinha Noh, Eunhee Kim, Hosu Lee et al.UbiComp 2026
- HearFit: Fitness Monitoring on Smart Speakers via Active Acoustic SensingYadong Xie, Fan Li, Yue Wu, Yu WangINFOCOM 2021 · 40 citations
- YouRefIt: Embodied Reference Understanding with Language and GestureYixin Chen, Qing Li, Deqian Kong, Yik Lun Kei et al.ICCV 2021 · 57 citations
