Pose Priors from Language Models
Sanjay Subramanian, Evonne Ng, Lea Müller, Dan Klein, Shiry Ginosar, Trevor Darrell
Abstract
Initial ProsePose Final Figure 1 . Optimizing human-to-human contacts in 3D pose. Our approach leverages the semantic priors of a Large Multimodal Model (LMM) to infer meaningful information about physical contact from images. Instead of relying on human annotations or motion capture data, we extract not only descriptive insights ("... engaged in a dance or embrace ...") but also structured constraints between body parts (underlined). By incorporating these LMM-derived constraints, we refine initial 3D human pose estimates, achieving realistic and semantically consistent reconstructions of contact. This scalable approach opens up new possibilities for contact-aware pose estimation without explicit contact annotations, making it a promising alternative to traditional methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f02e23c1-7552-4e25-a491-55ebc7eb5c30Cited by top-tier papers3
- PromptHMR: Promptable Human Mesh RecoveryYufu Wang, Yu Sun, Priyanka Patel, Kostas Daniilidis et al.CVPR 2025
- PI-HMR: Towards Robust In-bed Temporal Human Shape Reconstruction with Contact Pressure SensingZiyu Wu, Yufan Xiong, Mengting Niu, Fangting Xie et al.CVPR 2025
- Reconstructing Close Human Interaction with Appearance and Proxemics ReasoningBuzhen Huang, Chen Li, Chongyang Xu, Dongyue Lu et al.CVPR 2025
Builds on22
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 1,139 citations
- ViTPose: Simple Vision Transformer Baselines for Human Pose EstimationYufei Xu, Jing Zhang, Qiming Zhang, Dacheng TaoNeurIPS 2022 · 1,105 citations
Related papers
- InteractVLM: 3D Interaction Reasoning from 2D Foundational ModelsSai Kumar Dwivedi, Dimitrije Antic, Shashank Tripathi, Omid Taheri et al.CVPR 2025
- Learning Knowledge from Textual Descriptions for 3D Human Pose EstimationYi Wu, Jingtian Li, Shangfei Wang, Guoming Li et al.AAAI 2026
- Universal 3D Shape Matching via Coarse-to-Fine Language GuidanceQinfeng Xiao, Guofeng Mei, Bo Yang, Zhang Liying et al.CVPR 2026 · 1 citation
- Resolving 3D Human Pose Ambiguities With 3D Scene ConstraintsMohamed Hassan, Vasileios Choutas, Dimitrios Tzionas, Michael J. BlackICCV 2019 · 384 citations
- Automatic Human Scene Interaction through Contact Estimation and Motion AdaptationMingrui Zhang, Ming Chen, Yan Zhou, Li Chen et al.ACM MM 2023
