Lune

CVPR2026顶会

CLEP: Contrastive Language-Pose Pretraining

Sen Jia, Huayu Wang, Hsiang-Wei Huang, Zhaochong An, Jenq-Neng Hwang, Huaping Zhang, Lei Li

出版方
2026年份

摘要

Aligning natural language descriptions with precise 3D human poses remains a big challenge due to the scarcity of effective pose representation mechanisms and large-scale, semantically rich datasets. To overcome these limitations, we first introduce CLEP-2M, the largest 3D pose-language dataset to date, comprising two million high-quality 3D pose-language pairs. This dataset provides a 20-fold increase in scale and far richer semantic diversity than existing benchmarks. Second, we propose CLEP, a novel contrastive pretraining framework. The core of CLEP is Hier-Former, a hierarchical pose encoder specifically designed for language alignment. Its key innovation is a Cross-Scale Attention Fusion (CSAF) mechanism that dynamically integrates features from the joint, limb, and body levels. This enables CLEP to precisely align complex, multi-scale text descriptions with the pose representation. Extensive experimental evaluations on CLEP-2M and PoseScript demonstrate that our method consistently outperforms existing approaches across a range of downstream tasks. CLEP shows exceptional zero-shot generalization, achieving a 34.8 mRecall on the human-annotated PoseScript-H benchmark-a nearly 6-fold improvement from the baseline. Furthermore, CLEP demonstrates superior performance on pose generation and fine-grained pose editing. These results establish CLEP as a strong multimodal foundation model for humancentric understanding and generation tasks.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext f1c9d71f-e5fb-4aea-a2f7-22866ff3c8c8

它引用的顶会 Paper29

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖