CLEP: Contrastive Language-Pose Pretraining
Sen Jia, Huayu Wang, Hsiang-Wei Huang, Zhaochong An, Jenq-Neng Hwang, Huaping Zhang, Lei Li
Abstract
Aligning natural language descriptions with precise 3D human poses remains a big challenge due to the scarcity of effective pose representation mechanisms and large-scale, semantically rich datasets. To overcome these limitations, we first introduce CLEP-2M, the largest 3D pose-language dataset to date, comprising two million high-quality 3D pose-language pairs. This dataset provides a 20-fold increase in scale and far richer semantic diversity than existing benchmarks. Second, we propose CLEP, a novel contrastive pretraining framework. The core of CLEP is Hier-Former, a hierarchical pose encoder specifically designed for language alignment. Its key innovation is a Cross-Scale Attention Fusion (CSAF) mechanism that dynamically integrates features from the joint, limb, and body levels. This enables CLEP to precisely align complex, multi-scale text descriptions with the pose representation. Extensive experimental evaluations on CLEP-2M and PoseScript demonstrate that our method consistently outperforms existing approaches across a range of downstream tasks. CLEP shows exceptional zero-shot generalization, achieving a 34.8 mRecall on the human-annotated PoseScript-H benchmark-a nearly 6-fold improvement from the baseline. Furthermore, CLEP demonstrates superior performance on pose generation and fine-grained pose editing. These results establish CLEP as a strong multimodal foundation model for humancentric understanding and generation tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f1c9d71f-e5fb-4aea-a2f7-22866ff3c8c8Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
Related papers
- SensorLM: Learning the Language of Wearable SensorsYuwei Zhang, Kumar Ayush, Siyuan Qiao, A. Ali Heydari et al.NeurIPS 2025 · 75 citations
- FG-CLIP: Fine-Grained Visual and Textual AlignmentChunyu Xie, Bin Wang, Fanjing Kong, Jincheng Li et al.ICML 2025
- Overcoming the Pitfalls of Vision-Language Model for Image-Text RetrievalFeifei Zhang, Sijia Qu, Fan Shi, Changsheng XuACM MM 2024 · 12 citations
- RWKV-CLIP: A Robust Vision-Language Representation LearnerTiancheng Gu, Kaicheng Yang, Xiang An, Ziyong Feng et al.EMNLP 2024 · 11 citations
- CLIP-Hand3D: Exploiting 3D Hand Pose Estimation via Context-Aware PromptingShaoxiang Guo, Qing Cai, Lin Qi, Junyu DongACM MM 2023 · 10 citations
