Lune

CVPR2025Top-tier venue

Pose Priors from Language Models

Sanjay Subramanian, Evonne Ng, Lea Müller, Dan Klein, Shiry Ginosar, Trevor Darrell

2025Year
3Top-tier citations

Abstract

Initial ProsePose Final Figure 1 . Optimizing human-to-human contacts in 3D pose. Our approach leverages the semantic priors of a Large Multimodal Model (LMM) to infer meaningful information about physical contact from images. Instead of relying on human annotations or motion capture data, we extract not only descriptive insights ("... engaged in a dance or embrace ...") but also structured constraints between body parts (underlined). By incorporating these LMM-derived constraints, we refine initial 3D human pose estimates, achieving realistic and semantically consistent reconstructions of contact. This scalable approach opens up new possibilities for contact-aware pose estimation without explicit contact annotations, making it a promising alternative to traditional methods.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext f02e23c1-7552-4e25-a491-55ebc7eb5c30

Cited by top-tier papers3

Ask how each one uses it

Builds on22

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines