PoseFix: Correcting 3D Human Poses with Natural Language
Ginger Delmas, Philippe Weinzaepfel, Francesc Moreno-Noguer, Grégory Rogez
Abstract
Automatically producing instructions to modify one's posture could open the door to endless applications, such as personalized coaching and in-home physical therapy. Tackling the reverse problem (i.e., refining a 3D pose based on some natural language feedback) could help for assisted 3D character animation or robot teaching, for instance. Although a few recent works explore the connections between natural language and 3D human pose, none focus on describing 3D body pose differences. In this paper, we tackle the problem of correcting 3D human poses with natural language. To this end, we introduce the Pose-Fix dataset, which consists of several thousand paired 3D poses and their corresponding text feedback, that describe how the source pose needs to be modified to obtain the target pose. We demonstrate the potential of this dataset on two tasks: (1) text-based pose editing, that aims at generating corrected 3D body poses given a query pose and a text modifier; and (2) correctional text generation, where instructions are generated based on the differences between two body poses. The dataset and the code are available at https://europe.naverlabs.com/research/ computer-vision/posefix/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9991407b-aa7f-49a4-820b-4d9ebf6dfb77Cited by top-tier papers21
- Iterative Motion Editing with Natural LanguagePurvi Goel, Kuan-Chieh Wang, C. Karen Liu, Kayvon FatahalianSIGGRAPH 2024 · 22 citations
- HanDiffuser: Text-to-Image Generation with Realistic Hand AppearancesSupreeth Narasimhaswamy, Uttaran Bhattacharya, Xiang Chen, Ishita Dasgupta et al.CVPR 2024 · 17 citations
- FineMotion: A Dataset and Benchmark with Both Spatial and Temporal Annotation for Fine-Grained Motion Generation and EditingBizhu Wu, Jinheng Xie, Meidan Ding, Zhe Kong et al.ICCV 2025 · 11 citations
- PoseLLaVA: Pose Centric Multimodal LLM for Fine-Grained 3D Pose ManipulationDong Feng, Ping Guo, Encheng Peng, Mingmin Zhu et al.AAAI 2025 · 7 citations
- Learning Skill-Attributes for Transferable Assessment in VideoKumar Ashutosh, Kristen GraumanNeurIPS 2025 · 6 citations
Builds on26
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
Related papers
- FIXMYPOSE: Pose Correctional Captioning and RetrievalHyounghun Kim, Abhay Zala, Graham Burri, Mohit BansalAAAI 2021 · 22 citations
- CigTime: Corrective Instruction Generation Through Inverse Motion EditingQihang Fang, Chengcheng Tang, Bugra Tekin, Yanchao YangNeurIPS 2024 · 3 citations
- Cross-Axis Feature Fusion with Joint-Wise Motion Difference Prediction for Text-Based 3D Human Motion EditingGyojin Han, Junmo KimCVPR 2026 · 2 citations
- High-fidelity 3D Face Generation from Natural Language DescriptionsMenghua Wu, Hao Zhu, Linjia Huang, Yiyu Zhuang et al.CVPR 2023
- Autocompose: Automatic Generation of Pose Transition Descriptions for Composed Pose Retrieval Using Multimodal LLMsYi-Ting Shen, Sungmin Eum, Doheon Lee, Rohit Shete et al.ICCV 2025
