ExpertAF: Expert Actionable Feedback from Video
Kumar Ashutosh, Tushar Nagarajan, Georgios Pavlakos, Kris Kitani, Kristen Grauman
Abstract
Feedback is essential for learning a new skill or improving one's current skill-level. However, current methods for skill-assessment from video only provide scores or compare demonstrations, leaving the burden of knowing what to do differently on the user. We introduce a novel method to generate actionable feedback (AF) from video of a person doing a physical activity, such as basketball or soccer. Our method takes a video demonstration and its accompanying 3D body pose and generates (1) free-form expert commentary describing what the person is doing well and what they could improve, and (2) a visual expert demonstration that incorporates the required corrections. We show how to leverage Ego-Exo4D's [29] videos of skilled activity and expert commentary together with a strong language model to create a weakly-supervised training dataset for this task, and we devise a multimodal video-language model to infer coaching feedback. Our method is able to reason across multi-modal input combinations to output fullspectrum, actionable coaching-expert commentary, expert video retrieval, and expert pose generation-outperforming strong vision-language models on both established metrics and human preference studies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c9ba5bc3-20a0-4c34-8b6b-73d8c8e9cc92Cited by top-tier papers5
- Vid2Coach: Transforming How-To Videos into Task AssistantsMina Huh, Zihui Xue, Ujjaini Das, Kumar Ashutosh et al.UIST 2025 · 9 citations
- Learning Skill-Attributes for Transferable Assessment in VideoKumar Ashutosh, Kristen GraumanNeurIPS 2025 · 6 citations
- SkillSight: Efficient First-Person Skill Assessment with GazeChi Hsuan Wu, Kumar Ashutosh, Kristen GraumanCVPR 2026 · 4 citations
- Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person PerspectivesKristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani et al.CVPR 2024
- EgoLife: Towards Egocentric Life AssistantJingkang Yang, Shuai Liu, Hongming Guo, Yuhao Dong et al.CVPR 2025
Builds on35
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
- Learnable Triangulation of Human PoseKarim Iskakov, Egor Burkov, Victor S. Lempitsky, Yury MalkovICCV 2019 · 419 citations
- VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text UnderstandingHu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko et al.EMNLP 2021 · 399 citations
- Humans in 4D: Reconstructing and Tracking Humans with TransformersShubham Goel, Georgios Pavlakos, Jathushan Rajasegaran, Angjoo Kanazawa et al.ICCV 2023 · 390 citations
Related papers
- TechCoach: Towards Technical-Point-Aware Descriptive Action CoachingYuan-Ming Li, An-Lan Wang, Ling-An Zeng, Kun-Yu Lin et al.AAAI 2026
- SoleCoach: Sole Pressure and IMU-based MLLMs for Skill CoachingToshihiro Hirano, Hitoshi Yoshihara, Yichen Peng, Chen-Chieh Liao et al.CHI 2026 · 1 citation
- AgentCoach: LLM-Based Adaptive Coaching Feedback for Motor Skill LearningDizhi Ma, Jiakun Yu, Xinyi Wang, Xiyun Hu et al.CHI 2026 · 1 citation
- CoachMe: Decoding Sport Elements with a Reference-Based Coaching Instruction Generation ModelWei-Hsin Yeh, Yu-An Su, Chih-Ning Chen, Yi-Hsueh Lin et al.ACL 2025 · 1 citation
- Open-domain Video Commentary GenerationEdison Marrese-Taylor, Yumi Hamazono, Tatsuya Ishigaki, Goran Topic et al.EMNLP 2022 · 2 citations
