Learning Skill-Attributes for Transferable Assessment in Video
Kumar Ashutosh, Kristen Grauman
Abstract
Skill assessment from video entails rating the quality of a person's physical performance and explaining what could be done better. Today's models specialize for an individual sport, and suffer from the high cost and scarcity of expert-level supervision across the long tail of sports. Towards closing that gap, we explore transferable video representations for skill assessment. Our CROSSTRAINER approach discovers skill-attributes-such as balance, control, and hand positioning-whose meaning transcends the boundaries of any given sport, then trains a multimodal language model to generate actionable feedback for a novel video, e.g., "lift hands more to generate more power" as well as its proficiency level, e.g., early expert. We validate the new model on multiple datasets for both cross-sport (transfer) and intra-sport (in-domain) settings, where it achieves gains up to 60% relative to the state of the art. By abstracting out the shared behaviors indicative of human skill, the proposed video representation generalizes substantially better than an array of existing techniques, enriching today's multimodal large language models. Project
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 51c7c83f-977e-433a-8e28-84eb0aa7b7bfCited by top-tier papers1
Ask how each one uses itBuilds on43
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 1,550 citations
Related papers
- ExpertAF: Expert Actionable Feedback from VideoKumar Ashutosh, Tushar Nagarajan, Georgios Pavlakos, Kris Kitani et al.CVPR 2025
- SoleCoach: Sole Pressure and IMU-based MLLMs for Skill CoachingToshihiro Hirano, Hitoshi Yoshihara, Yichen Peng, Chen-Chieh Liao et al.CHI 2026 · 1 citation
- MoTrans: Customized Motion Transfer with Text-driven Video Diffusion ModelsXiaomin Li, Xu Jia, Qinghe Wang, Haiwen Diao et al.ACM MM 2024 · 6 citations
- VideoSetDiff: Identifying and Reasoning Similarities and Differences in Similar VideosYue Qiu, Yanjun Sun, Takuma Yagi, Shusaku Egami et al.ICCV 2025 · 1 citation
- SemTra: A Semantic Skill Translator for Cross-Domain Zero-Shot Policy AdaptationSangwoo Shin, Minjong Yoo, Jeongwoo Lee, Honguk WooAAAI 2024 · 6 citations
