Referring Human Pose and Mask Estimation In the Wild
Bo Miao, Mingtao Feng, Zijie Wu, Mohammed Bennamoun, Yongsheng Gao, Ajmal Mian
Abstract
We introduce Referring Human Pose and Mask Estimation (R-HPM) in the wild, where either a text or positional prompt specifies the person of interest in an image. This new task holds significant potential for human-centric applications such as assistive robotics and sports analysis. In contrast to previous works, R-HPM (i) ensures high-quality, identity-aware results corresponding to the referred person, and (ii) simultaneously predicts human pose and mask for a comprehensive representation. To achieve this, we introduce a large-scale dataset named RefHuman, which substantially extends the MS COCO dataset with additional text and positional prompt annotations. RefHuman includes over 50,000 annotated instances in the wild, each equipped with keypoint, mask, and prompt annotations. To enable prompt-conditioned estimation, we propose the first end-to-end promptable approach named UniPHD for R-HPM. UniPHD extracts multimodal representations and employs a proposed pose-centric hierarchical decoder to process (text or positional) instance queries and keypoint queries, producing results specific to the referred person. Extensive experiments demonstrate that UniPHD produces quality results based on user-friendly prompts and achieves top-tier performance on RefHuman val and MS COCO val2017. Data and Code: https://github.com/bo-miao/RefHuman
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 060f37ce-e578-4426-8420-d32eaf6a1cdbCited by top-tier papers2
- CymbaDiff: Structured Spatial Diffusion for Sketch-based 3D Semantic Urban Scene GenerationLi Liang, Bo Miao, Xinyu Wang, Naveed Akhtar et al.NeurIPS 2025 · 4 citations
- Disentangled Hierarchical VAE for 3D Human-Human Interaction GenerationZichen Geng, Zeeshan Hayder, Bo Miao, Jian Liu et al.ICLR 2026 · 3 citations
Builds on33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve et al.ICCV 2021 · 1,114 citations
Related papers
- PromptHMR: Promptable Human Mesh RecoveryYufu Wang, Yu Sun, Priyanka Patel, Kostas Daniilidis et al.CVPR 2025
- SAM 3D Body: Robust Full-Body Human Mesh RecoveryXitong Yang, Devansh Kukreja, Don Pinkus, Taosha Fan et al.CVPR 2026 · 81 citations
- FinePOSE: Fine-Grained Prompt-Driven 3D Human Pose Estimation via Diffusion ModelsJinglin Xu, Yijie Guo, Yuxin PengCVPR 2024 · 39 citations
- Gaze Target Estimation Anywhere with ConceptsXu Cao, Houze Yang, Vipin Gunda, Zhongyi Zhou et al.CVPR 2026 · 3 citations
- Referring to Any PersonQing Jiang, Lin Wu, Zhaoyang Zeng, Tianhe Ren et al.ICCV 2025 · 1 citation
