CiCo: Domain-Aware Sign Language Retrieval via Cross-Lingual Contrastive Learning
Yiting Cheng, Fangyun Wei, Jianmin Bao, Dong Chen, Wenqiang Zhang
Abstract
This work focuses on sign language retrieval-a recently proposed task for sign language understanding. Sign language retrieval consists of two sub-tasks: text-to-sign-video (T2V) retrieval and sign-video-to-text (V2T) retrieval. Different from traditional video-text retrieval, sign language videos, not only contain visual signals but also carry abundant semantic meanings by themselves due to the fact that sign languages are also natural languages. Considering this character, we formulate sign language retrieval as a cross-lingual retrieval problem as well as a video-text retrieval task. Concretely, we take into account the linguistic properties of both sign languages and natural languages, and simultaneously identify the fine-grained cross-lingual (i.e., sign-to-word) mappings while contrasting the texts and the sign videos in a joint embedding space. This process is termed as cross-lingual contrastive learning. Another challenge is raised by the data scarcity issue-sign language datasets are orders of magnitude smaller in scale than that of speech recognition. We alleviate this issue by adopting a domain-agnostic sign encoder pre-trained on large-scale sign videos into the target domain via pseudolabeling. Our framework, termed as domain-aware sign language retrieval via Cross-lingual Contrastive learning or CiCo for short, outperforms the pioneering method by large margins on various datasets, e.g., +22.4 T2V and +28.0 V2T R@1 improvements on How2Sign dataset, and +13.7 T2V and +17.1 V2T R@1 improvements on PHOENIX-2014T dataset. Code and models are available at: https://github.com/FangyunWei/SLRT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0e9ccc59-42fd-4f33-90e7-918a6eb81f79Cited by top-tier papers11
- Unified Coarse-to-Fine Alignment for Video-Text RetrievalZiyang Wang, Yi-Lin Sung, Feng Cheng, Gedas Bertasius et al.ICCV 2023 · 90 citations
- Improving Gloss-free Sign Language Translation by Reducing Representation DensityJinhui Ye, Xing Wang, Wenxiang Jiao, Junwei Liang et al.NeurIPS 2024 · 49 citations
- Improving Continuous Sign Language Recognition with Cross-Lingual SignsFangyun Wei, Yutong ChenICCV 2023 · 46 citations
- Contextual Augmented Global Contrast for Multimodal Intent RecognitionKaili Sun, Zhiwen Xie, Mang Ye, Huyin ZhangCVPR 2024 · 19 citations
- Geo-Sign: Hyperbolic Contrastive Regularisation for Geometrically Aware Sign Language TranslationEdward Fish, Richard BowdenNeurIPS 2025 · 15 citations
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 1,550 citations
Related papers
- SignCLIP: Connecting Text and Sign Language by Contrastive LearningZifan Jiang, Gerard Sant, Amit Moryossef, Mathias Müller et al.EMNLP 2024 · 4 citations
- Sign Language Video Retrieval with Free-Form Textual QueriesAmanda Cardoso Duarte, Samuel Albanie, Xavier Giró-i-Nieto, Gül VarolCVPR 2022 · 27 citations
- CVT-SLR: Contrastive Visual-Textual Transformation for Sign Language Recognition with Variational AlignmentJiangbin Zheng, Yile Wang, Cheng Tan, Siyuan Li et al.CVPR 2023
- Contrastive Disentangled Meta-Learning for Signer-Independent Sign Language TranslationTao Jin, Zhou ZhaoACM MM 2021 · 23 citations
- Semantic Hardness Is Not Visual Hardness: Sign-Aware Hard Negative Mining for Sign Language RetrievalJunmyeong Lee, Chan Hur, ChangSu Choi, Sukmin Cho et al.ACL 2026
