SignCLIP: Connecting Text and Sign Language by Contrastive Learning
Zifan Jiang, Gerard Sant, Amit Moryossef, Mathias Müller, Rico Sennrich, Sarah Ebling
Abstract
We present SignCLIP, which re-purposes CLIP (Contrastive Language-Image Pretraining) to project spoken language text and sign language videos, two classes of natural languages of distinct modalities, into the same space. SignCLIP is an efficient method of learning useful visual representations for sign language processing from large-scale, multilingual video-text pairs, without directly optimizing for a specific task or sign language which is often of limited size. We pretrain SignCLIP on Spreadthesign, a prominent sign language dictionary consisting of ∼500 thousand video clips in up to 44 sign languages, and evaluate it with various downstream datasets. SignCLIP discerns in-domain signing with notable text-to-video/video-to-text retrieval accuracy. It also performs competitively for out-of-domain downstream tasks such as isolated sign language recognition upon essential few-shot prompting or fine-tuning. We analyze the latent space formed by the spoken language text and sign language poses, which provides additional linguistic insights. Our code and models are openly available 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Stable Signer: Hierarchical Sign Language Generative ModelSen Fang, Yalin Feng, Hongbin Zhong, Yanxin Zhang et al.ACL 2026 · 3 citations
- Logos as a Well-Tempered Pre-train for Sign Language RecognitionIlya Ovodov, Petr Surovtsev, Karina Kvanchiani, Alexander Kapitanov et al.EMNLP 2025 · 1 citation
- SHuBERT: Self-Supervised Sign Language Representation Learning via Multi-Stream Cluster PredictionShester Gueuwou, Xiaodan Du, Greg Shakhnarovich, Karen Livescu et al.ACL 2025
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei et al.CVPR 2022 · 1,847 citations
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
- VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text UnderstandingHu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko et al.EMNLP 2021 · 399 citations
Related papers
- CiCo: Domain-Aware Sign Language Retrieval via Cross-Lingual Contrastive LearningYiting Cheng, Fangyun Wei, Jianmin Bao, Dong Chen et al.CVPR 2023
- RankCLIP: Ranking-Consistent Language-Image PretrainingYiming Zhang, Zhuokai Zhao, Zhaorun Chen, Zhili Feng et al.ICCV 2025 · 1 citation
- How Much Can CLIP Benefit Vision-and-Language Tasks?Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal et al.ICLR 2022 · 503 citations
- SpaceCLIP: A Vision-Language Pretraining Framework With Spatial Reconstruction On TextBo Zou, Chao Yang, Chengbin Quan, Youjian ZhaoACM MM 2023 · 1 citation
- DiffCLIP: Few-shot Language-driven Multimodal ClassifierJiaqing Zhang, Mingxiang Cao, Xue Yang, Kai Jiang et al.AAAI 2025 · 4 citations
