FIXMYPOSE: Pose Correctional Captioning and Retrieval
Hyounghun Kim, Abhay Zala, Graham Burri, Mohit Bansal
摘要
Interest in physical therapy and individual exercises such as yoga/dance has increased alongside the well-being trend, and people globally enjoy such exercises at home/office via video streaming platforms. However, such exercises are hard to follow without expert guidance. Even if experts can help, it is almost impossible to give personalized feedback to every trainee remotely. Thus, automated pose correction systems are required more than ever, and we introduce a new captioning dataset named FixMyPose to address this need. We collect natural language descriptions of correcting a “current” pose to look like a “target” pose. To support a multilingual setup, we collect descriptions in both English and Hindi. The collected descriptions have interesting linguistic properties such as egocentric relations to the environment objects, analogous references, etc., requiring an understanding of spatial relations and commonsense knowledge about postures. Further, to avoid ML biases, we maintain a balance across characters with diverse demographics, who perform a variety of movements in several interior environments (e.g., homes, offices). From our FixMyPose dataset, we introduce two tasks: the pose-correctional-captioning task and its reverse, the target-pose-retrieval task. During the correctional-captioning task, models must generate the descriptions of how to move from the current to the target pose image, whereas in the retrieval task, models should select the correct target pose given the initial pose and the correctional description. We present strong cross-attention baseline models (uni/multimodal, RL, multilingual) and also show that our baselines are competitive with other models when evaluated on other image-difference datasets. We also propose new task-specific metrics (object-match, body-part-match, direction-match) and conduct human evaluation for more reliable evaluation, and we demonstrate a large human-model performance gap suggesting room for promising future work. Finally, to verify the sim-to-real transfer of our FixMyPose dataset, we collect a set of real images and show promising performance on these images. Data and code are available: https://fixmypose-unc.github.io.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- PoseFix: Correcting 3D Human Poses with Natural LanguageGinger Delmas, Philippe Weinzaepfel, Francesc Moreno-Noguer, Grégory RogezICCV 2023 · 被引用 49 次
- ProtoRes: Proto-Residual Network for Pose Authoring via Learned Inverse KinematicsBoris N. Oreshkin, Florent Bocquelet, Félix G. Harvey, Bay Raitt 等ICLR 2022 · 被引用 18 次
- CigTime: Corrective Instruction Generation Through Inverse Motion EditingQihang Fang, Chengcheng Tang, Bugra Tekin, Yanchao YangNeurIPS 2024 · 被引用 3 次
- UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and EditingYiheng Li, Ruibing Hou, Hong Chang, Shiguang Shan 等CVPR 2025
它引用的顶会 Paper5
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearchXin Wang, Jiawei Wu, Jun-Kun Chen, Lei Li 等ICCV 2019 · 被引用 688 次
- Robust Change CaptioningDong Huk Park, Trevor Darrell, Anna RohrbachICCV 2019 · 被引用 217 次
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 被引用 40 次
- Reducing Task Load with an Embodied Intelligent Virtual Assistant for Improved Performance in Collaborative Decision MakingKangsoo Kim, Celso M. de Melo, Nahal Norouzi, Gerd Bruder 等IEEE VR 2020 · 被引用 22 次
相关 Paper
- Less is More: Generating Grounded Navigation Instructions from LandmarksSu Wang, Ceslee Montgomery, Jordi Orbay, Vighnesh Birodkar 等CVPR 2022 · 被引用 41 次
- Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal GroundingAlexander Ku, Peter Anderson, Roma Patel, Eugene Ie 等EMNLP 2020 · 被引用 208 次
- YouRefIt: Embodied Reference Understanding with Language and GestureYixin Chen, Qing Li, Deqian Kong, Yik Lun Kei 等ICCV 2021 · 被引用 57 次
- Autocompose: Automatic Generation of Pose Transition Descriptions for Composed Pose Retrieval Using Multimodal LLMsYi-Ting Shen, Sungmin Eum, Doheon Lee, Rohit Shete 等ICCV 2025
- OmniDiff: A Comprehensive Benchmark for Fine-Grained Image Difference CaptioningYuan Liu, Saihui Hou, Saijie Hou, Jiabao Du 等ICCV 2025 · 被引用 1 次
