CLIP-Hand3D: Exploiting 3D Hand Pose Estimation via Context-Aware Prompting
Shaoxiang Guo, Qing Cai, Lin Qi, Junyu Dong
摘要
Contrastive Language-Image Pre-training (CLIP) starts to emerge in many computer vision tasks and has achieved promising performance. However, it remains underexplored whether CLIP can be generalized to 3D hand pose estimation, as bridging text prompts with pose-aware features presents significant challenges due to the discrete nature of joint positions in 3D space. In this paper, we make one of the first attempts to propose a novel 3D hand pose estimator from monocular images, dubbed as CLIP-Hand3D, which successfully bridges the gap between text prompts and irregular detailed pose distribution. In particular, the distribution order of hand joints in various 3D space directions is derived from pose labels, forming corresponding text prompts that are subsequently encoded into text representations. Simultaneously, 21 hand joints in the 3D space are retrieved, and their spatial distribution (in x, y, and z axes) is encoded to form pose-aware features. Subsequently, we maximize semantic consistency for a pair of pose-text features following a CLIP-based contrastive learning paradigm. Furthermore, a coarse-to-fine mesh regressor is designed, which is capable of effectively querying joint-aware cues from the feature pyramid. Extensive experiments on several public hand benchmarks show that the proposed model attains a significantly faster inference speed while achieving state-of-the-art performance compared to methods utilizing the similar scale backbone. Code is available at: https://anonymous.4open.science/r/CLIP_Hand_Demo-FD2B/ README.md.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Differential Contrastive Training for Gaze EstimationLin Zhang, Yi Tian, Xiyun Wang, Wanru Xu 等ACM MM 2025 · 被引用 5 次
- Seeing Beyond the Crop: Using Language Priors for Out-of-Bounding Box Keypoint PredictionBavesh Balaji, Jerrin Bright, Yuhao Chen, Sirisha Rambhatla 等NeurIPS 2024 · 被引用 4 次
- SGTC: Semantic-Guided Triplet Co-training for Sparsely Annotated Semi-Supervised Medical Image SegmentationKe Yan, Qing Cai, Fan Zhang, Ziyan Cao 等AAAI 2025 · 被引用 1 次
它引用的顶会 Paper33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- DenseCLIP: Language-Guided Dense Prediction with Context-Aware PromptingYongming Rao, Wenliang Zhao, Guangyi Chen, Yansong Tang 等CVPR 2022 · 被引用 527 次
- Exploiting Spatial-Temporal Relationships for 3D Pose Estimation via Graph Convolutional NetworksYujun Cai, Liuhao Ge, Jun Liu, Jianfei Cai 等ICCV 2019 · 被引用 504 次
- FreiHAND: A Dataset for Markerless Capture of Hand Pose and Shape From Single RGB ImagesChristian Zimmermann, Duygu Ceylan, Jimei Yang, Bryan C. Russell 等ICCV 2019 · 被引用 493 次
- Mesh GraphormerKevin Lin, Lijuan Wang, Zicheng LiuICCV 2021 · 被引用 399 次
相关 Paper
- CLAMP: Prompt-based Contrastive Learning for Connecting Language and Animal PoseXu Zhang, Wen Wang, Zhe Chen, Yufei Xu 等CVPR 2023
- CrowdCLIP: Unsupervised Crowd Counting via Vision-Language ModelDingkang Liang, Jiahao Xie, Zhikang Zou, Xiaoqing Ye 等CVPR 2023
- CLIP-6D: Empowering CLIP as a Zero-Shot 6D Pose Estimator Through Generalizable Object-Specific RepresentationsHua Wang, Hong Liu, Jiale Ren, Mingxin Tan 等ACM MM 2025 · 被引用 1 次
- CLIP2UDA: Making Frozen CLIP Reward Unsupervised Domain Adaptation in 3D Semantic SegmentationYao Wu, Mingwei Xing, Yachao Zhang, Yuan Xie 等ACM MM 2024 · 被引用 12 次
- HairCLIP: Design Your Hair by Text and Reference ImageTianyi Wei, Dongdong Chen, Wenbo Zhou, Jing Liao 等CVPR 2022 · 被引用 94 次
