Gloss Attention for Gloss-free Sign Language Translation
Aoxiong Yin, Tianyun Zhong, Li Tang, Weike Jin, Tao Jin, Zhou Zhao
Abstract
Most sign language translation (SLT) methods to date require the use of gloss annotations to provide additional supervision information, however, the acquisition of gloss is not easy. To solve this problem, we first perform an analysis of existing models to confirm how gloss annotations make SLT easier. We find that it can provide two aspects of information for the model, 1) it can help the model implicitly learn the location of semantic boundaries in continuous sign language videos, 2) it can help the model understand the sign language video globally. We then propose gloss attention, which enables the model to keep its attention within video segments that have the same semantics locally, just as gloss helps existing models do. Furthermore, we transfer the knowledge of sentence-to-sentence similarity from the natural language model to our gloss attention SLT network (GASLT) to help it understand sign language videos at the sentence level. Experimental results on multiple large-scale sign language datasets show that our proposed GASLT model significantly outperforms existing methods. Our code is provided in https://github . com/YinAoXiong/GASLT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 59381c98-9e96-449a-bc02-b70818e10502Cited by top-tier papers27
- Gloss-free Sign Language Translation: Improving from Visual-Language PretrainingBenjia Zhou, Zhigang Chen, Albert Clapés, Jun Wan et al.ICCV 2023 · 123 citations
- Sign2GPT: Leveraging Large Language Models for Gloss-Free Sign Language TranslationRyan Wong, Necati Cihan Camgöz, Richard BowdenICLR 2024 · 58 citations
- Improving Gloss-free Sign Language Translation by Reducing Representation DensityJinhui Ye, Xing Wang, Wenxiang Jiao, Junwei Liang et al.NeurIPS 2024 · 49 citations
- Bridging Sign and Spoken Languages: Pseudo Gloss Generation for Sign Language TranslationJianyuan Guo, Peike Li, Trevor CohnNeurIPS 2025 · 17 citations
- TranSpeech: Speech-to-Speech Translation With Bilateral PerturbationRongjie Huang, Jinglin Liu, Huadai Liu, Yi Ren et al.ICLR 2023 · 17 citations
Builds on22
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Two-Stream Network for Sign Language Recognition and TranslationYutong Chen, Ronglai Zuo, Fangyun Wei, Yu Wu et al.NeurIPS 2022 · 288 citations
- Spatial-Temporal Multi-Cue Network for Continuous Sign Language RecognitionHao Zhou, Wengang Zhou, Yun Zhou, Houqiang LiAAAI 2020 · 249 citations
- Visual Alignment Constraint for Continuous Sign Language RecognitionYuecong Min, Aiming Hao, Xiujuan Chai, Xilin ChenICCV 2021 · 211 citations
- ProDiff: Progressive Fast Diffusion Model for High-Quality Text-to-SpeechRongjie Huang, Zhou Zhao, Huadai Liu, Jinglin Liu et al.ACM MM 2022 · 182 citations
Related papers
- Gloss-Free End-to-End Sign Language TranslationKezhou Lin, Xiaohan Wang, Linchao Zhu, Ke Sun et al.ACL 2023 · 25 citations
- Sentence-level Segmentation for Long Sign Language Videos with CaptionsBowen Guo, Shiwei Gan, Yafeng Yin, Xiao Liu et al.ACM MM 2025
- Semi-Supervised Spoken Language GlossificationHuijie Yao, Wengang Zhou, Hao Zhou, Houqiang LiACL 2024
- Learning Effective Sign Features without Text for Gloss-free Sign Language TranslationShiwei Gan, Xiao Liu, Yafeng Yin, Nan Liu et al.CVPR 2026 · 2 citations
- BoostSLT: Boosting Sign Language Translation via a Plug-and-Play Diffusion-Based Semantic EnhancerChangzhou Han, Wanlun Ma, Xi Tang, Kun Hu et al.CVPR 2026
