Pose Priors from Language Models
Sanjay Subramanian, Evonne Ng, Lea Müller, Dan Klein, Shiry Ginosar, Trevor Darrell
摘要
Initial ProsePose Final Figure 1 . Optimizing human-to-human contacts in 3D pose. Our approach leverages the semantic priors of a Large Multimodal Model (LMM) to infer meaningful information about physical contact from images. Instead of relying on human annotations or motion capture data, we extract not only descriptive insights ("... engaged in a dance or embrace ...") but also structured constraints between body parts (underlined). By incorporating these LMM-derived constraints, we refine initial 3D human pose estimates, achieving realistic and semantically consistent reconstructions of contact. This scalable approach opens up new possibilities for contact-aware pose estimation without explicit contact annotations, making it a promising alternative to traditional methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- PromptHMR: Promptable Human Mesh RecoveryYufu Wang, Yu Sun, Priyanka Patel, Kostas Daniilidis 等CVPR 2025
- PI-HMR: Towards Robust In-bed Temporal Human Shape Reconstruction with Contact Pressure SensingZiyu Wu, Yufan Xiong, Mengting Niu, Fangting Xie 等CVPR 2025
- Reconstructing Close Human Interaction with Appearance and Proxemics ReasoningBuzhen Huang, Chen Li, Chongyang Xu, Dongyue Lu 等CVPR 2025
它引用的顶会 Paper22
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 被引用 1,139 次
- ViTPose: Simple Vision Transformer Baselines for Human Pose EstimationYufei Xu, Jing Zhang, Qiming Zhang, Dacheng TaoNeurIPS 2022 · 被引用 1,105 次
相关 Paper
- InteractVLM: 3D Interaction Reasoning from 2D Foundational ModelsSai Kumar Dwivedi, Dimitrije Antic, Shashank Tripathi, Omid Taheri 等CVPR 2025
- Learning Knowledge from Textual Descriptions for 3D Human Pose EstimationYi Wu, Jingtian Li, Shangfei Wang, Guoming Li 等AAAI 2026
- Universal 3D Shape Matching via Coarse-to-Fine Language GuidanceQinfeng Xiao, Guofeng Mei, Bo Yang, Zhang Liying 等CVPR 2026 · 被引用 1 次
- Resolving 3D Human Pose Ambiguities With 3D Scene ConstraintsMohamed Hassan, Vasileios Choutas, Dimitrios Tzionas, Michael J. BlackICCV 2019 · 被引用 384 次
- Automatic Human Scene Interaction through Contact Estimation and Motion AdaptationMingrui Zhang, Ming Chen, Yan Zhou, Li Chen 等ACM MM 2023
