From 3D Pose to Prose: Biomechanics-Grounded Vision-Language Coaching
Yuyang Ji, Yixuan Shen, Shengjie Zhu, Yu Kong, Feng Liu
Abstract
We present BioCoach, a biomechanics-grounded visionlanguage framework for fitness coaching from streaming video. BioCoach fuses visual appearance and 3D skeletal kinematics, through a novel three-stage pipeline: an exercise-specific degree-of-freedom selector that focuses analysis on salient joints; a structured biomechanical context that pairs individualized morphometrics with cycle and constraint analysis; and a vision-biomechanics conditioned feedback module that applies cross-attention to generate precise, actionable text. Using parameter-efficient training that freezes the vision and language backbones, BioCoach yields transparent, personalized reasoning rather than pattern matching. To enable learning and fair evaluation, we augment QEVD-fit-coach with biomechanicsoriented feedback to create QEVD-bio-fit-coach, and we introduce a biomechanics-aware LLM judge metric. Bio-Coach delivers clear gains on QEVD-bio-fit-coach across lexical and judgment metrics while maintaining temporal triggering; on the original QEVD-fit-coach, it improves text quality and correctness with near-parity timing, demonstrating that explicit kinematics and constraints are key to accurate, phase-aware coaching. Project
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8d4fe99f-544e-438c-9dd8-31cf791b7f5dBuilds on19
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- MotionGPT: Human Motion as a Foreign LanguageBiao Jiang, Xin Chen, Wen Liu, Jingyi Yu et al.NeurIPS 2023 · 698 citations
Related papers
- AgentCoach: LLM-Based Adaptive Coaching Feedback for Motor Skill LearningDizhi Ma, Jiakun Yu, Xinyi Wang, Xiyun Hu et al.CHI 2026 · 1 citation
- TechCoach: Towards Technical-Point-Aware Descriptive Action CoachingYuan-Ming Li, An-Lan Wang, Ling-An Zeng, Kun-Yu Lin et al.AAAI 2026
- ExpertAF: Expert Actionable Feedback from VideoKumar Ashutosh, Tushar Nagarajan, Georgios Pavlakos, Kris Kitani et al.CVPR 2025
- ViSTAR: Virtual Skill Training with Augmented Reality with 3D Avatars and LLM coaching agentChunggi Lee, Hayato Saiki, Tica Lin, Eiji Ikeda et al.CHI 2026 · 1 citation
- VisMimic: Integrating Motion Chain in Feedback Video Generation for Motor CoachingLiqi Cheng, Xiao Xie, Yiwei Peng, Minghao Feng et al.UIST 2025 · 2 citations
