CLIP-Hand3D: Exploiting 3D Hand Pose Estimation via Context-Aware Prompting
Shaoxiang Guo, Qing Cai, Lin Qi, Junyu Dong
Abstract
Contrastive Language-Image Pre-training (CLIP) starts to emerge in many computer vision tasks and has achieved promising performance. However, it remains underexplored whether CLIP can be generalized to 3D hand pose estimation, as bridging text prompts with pose-aware features presents significant challenges due to the discrete nature of joint positions in 3D space. In this paper, we make one of the first attempts to propose a novel 3D hand pose estimator from monocular images, dubbed as CLIP-Hand3D, which successfully bridges the gap between text prompts and irregular detailed pose distribution. In particular, the distribution order of hand joints in various 3D space directions is derived from pose labels, forming corresponding text prompts that are subsequently encoded into text representations. Simultaneously, 21 hand joints in the 3D space are retrieved, and their spatial distribution (in x, y, and z axes) is encoded to form pose-aware features. Subsequently, we maximize semantic consistency for a pair of pose-text features following a CLIP-based contrastive learning paradigm. Furthermore, a coarse-to-fine mesh regressor is designed, which is capable of effectively querying joint-aware cues from the feature pyramid. Extensive experiments on several public hand benchmarks show that the proposed model attains a significantly faster inference speed while achieving state-of-the-art performance compared to methods utilizing the similar scale backbone. Code is available at: https://anonymous.4open.science/r/CLIP_Hand_Demo-FD2B/ README.md.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 11e2d52b-abec-477f-9413-7fafa8d1d9d9Cited by top-tier papers3
- Differential Contrastive Training for Gaze EstimationLin Zhang, Yi Tian, Xiyun Wang, Wanru Xu et al.ACM MM 2025 · 5 citations
- Seeing Beyond the Crop: Using Language Priors for Out-of-Bounding Box Keypoint PredictionBavesh Balaji, Jerrin Bright, Yuhao Chen, Sirisha Rambhatla et al.NeurIPS 2024 · 4 citations
- SGTC: Semantic-Guided Triplet Co-training for Sparsely Annotated Semi-Supervised Medical Image SegmentationKe Yan, Qing Cai, Fan Zhang, Ziyan Cao et al.AAAI 2025 · 1 citation
Builds on33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- DenseCLIP: Language-Guided Dense Prediction with Context-Aware PromptingYongming Rao, Wenliang Zhao, Guangyi Chen, Yansong Tang et al.CVPR 2022 · 527 citations
- Exploiting Spatial-Temporal Relationships for 3D Pose Estimation via Graph Convolutional NetworksYujun Cai, Liuhao Ge, Jun Liu, Jianfei Cai et al.ICCV 2019 · 504 citations
- FreiHAND: A Dataset for Markerless Capture of Hand Pose and Shape From Single RGB ImagesChristian Zimmermann, Duygu Ceylan, Jimei Yang, Bryan C. Russell et al.ICCV 2019 · 493 citations
- Mesh GraphormerKevin Lin, Lijuan Wang, Zicheng LiuICCV 2021 · 399 citations
Related papers
- CLAMP: Prompt-based Contrastive Learning for Connecting Language and Animal PoseXu Zhang, Wen Wang, Zhe Chen, Yufei Xu et al.CVPR 2023
- CrowdCLIP: Unsupervised Crowd Counting via Vision-Language ModelDingkang Liang, Jiahao Xie, Zhikang Zou, Xiaoqing Ye et al.CVPR 2023
- CLIP-6D: Empowering CLIP as a Zero-Shot 6D Pose Estimator Through Generalizable Object-Specific RepresentationsHua Wang, Hong Liu, Jiale Ren, Mingxin Tan et al.ACM MM 2025 · 1 citation
- CLIP2UDA: Making Frozen CLIP Reward Unsupervised Domain Adaptation in 3D Semantic SegmentationYao Wu, Mingwei Xing, Yachao Zhang, Yuan Xie et al.ACM MM 2024 · 12 citations
- HairCLIP: Design Your Hair by Text and Reference ImageTianyi Wei, Dongdong Chen, Wenbo Zhou, Jing Liao et al.CVPR 2022 · 94 citations
