Bridging Sign and Spoken Languages: Pseudo Gloss Generation for Sign Language Translation
Jianyuan Guo, Peike Li, Trevor Cohn
Abstract
Sign Language Translation (SLT) aims to map sign language videos to spoken language text. A common approach relies on gloss annotations as an intermediate representation, decomposing SLT into two sub-tasks: video-to-gloss recognition and gloss-to-text translation. While effective, this paradigm depends on expert-annotated gloss labels, which are costly and rarely available in existing datasets, limiting its scalability. To address this challenge, we propose a gloss-free pseudo gloss generation framework that eliminates the need for human-annotated glosses while preserving the structured intermediate representation. Specifically, we prompt a Large Language Model (LLM) with a few example text-gloss pairs using in-context learning to produce draft sign glosses from spoken language text. To enhance the correspondence between LLM-generated pseudo glosses and the sign sequences in video, we correct the ordering in the pseudo glosses for better alignment via a weakly supervised learning process. This reordering facilitates the incorporation of auxiliary alignment objectives, and allows for the use of efficient supervision via a Connectionist Temporal Classification (CTC) loss. We train our SLT mode, which consists of a vision encoder and a translator, through a three-stage pipeline, which progressively narrows the modality gap between sign language and spoken language. Despite its simplicity, our approach outperforms previous state-of-the-art gloss-free frameworks on two SLT benchmarks and achieves competitive results compared to gloss-based methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4fef1d2e-7215-45ab-b042-edbc4f451dcaCited by top-tier papers2
- Learning Effective Sign Features without Text for Gloss-free Sign Language TranslationShiwei Gan, Xiao Liu, Yafeng Yin, Nan Liu et al.CVPR 2026 · 2 citations
- CNSL-bench: Benchmarking the Sign Language Understanding Capabilities of MLLMs on Chinese National Sign LanguageRui Zhao, Xuewen Zhong, Xiaoyun Zheng, Jinsong Su et al.ACL 2026
Builds on20
- Two-Stream Network for Sign Language Recognition and TranslationYutong Chen, Ronglai Zuo, Fangyun Wei, Yu Wu et al.NeurIPS 2022 · 288 citations
- TSPNet: Hierarchical Feature Learning via Temporal Semantic Pyramid for Sign Language TranslationDongxu Li, Chenchen Xu, Xin Yu, Kaihao Zhang et al.NeurIPS 2020 · 171 citations
- A Simple Multi-Modality Transfer Learning Baseline for Sign Language TranslationYutong Chen, Fangyun Wei, Xiao Sun, Zhirong Wu et al.CVPR 2022 · 137 citations
- SignBERT: Pre-Training of Hand-Model-Aware Representation for Sign Language RecognitionHezhen Hu, Weichao Zhao, Wengang Zhou, Yuechen Wang et al.ICCV 2021 · 125 citations
- Gloss-free Sign Language Translation: Improving from Visual-Language PretrainingBenjia Zhou, Zhigang Chen, Albert Clapés, Jun Wan et al.ICCV 2023 · 123 citations
Related papers
- Leveraging the Power of MLLMs for Gloss-Free Sign Language TranslationJungeun Kim, Hyeongwoo Jeon, Jongseong Bae, Ha Young KimICCV 2025 · 10 citations
- C2ST: Cross-modal Contextualized Sequence Transduction for Continuous Sign Language RecognitionHuaiwen Zhang, Zihang Guo, Yang Yang, Xin Liu et al.ICCV 2023 · 21 citations
- LLMs are Good Sign Language TranslatorsJia Gong, Lin Geng Foo, Yixuan He, Hossein Rahmani et al.CVPR 2024
- Gloss-Free End-to-End Sign Language TranslationKezhou Lin, Xiaohan Wang, Linchao Zhu, Ke Sun et al.ACL 2023 · 25 citations
- Sign2GPT: Leveraging Large Language Models for Gloss-Free Sign Language TranslationRyan Wong, Necati Cihan Camgöz, Richard BowdenICLR 2024 · 58 citations
