Open-Domain Sign Language Translation Learned from Online Video
Bowen Shi, Diane Brentari, Gregory Shakhnarovich, Karen Livescu
Abstract
Existing work on sign language translation – that is, translation from sign language videos into sentences in a written language – has focused mainly on (1) data collected in a controlled environment or (2) data in a specific domain, which limits the applicability to real-world settings. In this paper, we introduce OpenASL, a large-scale American Sign Language (ASL) - English dataset collected from online video sites (e.g., YouTube).OpenASL contains 288 hours of ASL videos in multiple domains from over 200 signers and is the largest publicly available ASL translation dataset to date. To tackle the challenges of sign language translation in realistic settings and without glosses, we propose a set of techniques including sign search as a pretext task for pre-training and fusion of mouthing and handshape features. The proposed techniques produce consistent and large improvements in translation quality, over baseline models basedon prior work.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dc579e4b-a9ce-40c3-ad93-c4537b40eb23Cited by top-tier papers20
- Sign2GPT: Leveraging Large Language Models for Gloss-Free Sign Language TranslationRyan Wong, Necati Cihan Camgöz, Richard BowdenICLR 2024 · 58 citations
- Gloss-Free End-to-End Sign Language TranslationKezhou Lin, Xiaohan Wang, Linchao Zhu, Ke Sun et al.ACL 2023 · 25 citations
- Scaling Sign Language TranslationBiao Zhang, Garrett Tanzer, Orhan FiratNeurIPS 2024 · 21 citations
- Bridging Sign and Spoken Languages: Pseudo Gloss Generation for Sign Language TranslationJianyuan Guo, Peike Li, Trevor CohnNeurIPS 2025 · 17 citations
- Towards AI-driven Sign Language Generation with Non-manual MarkersHan Zhang, Rotem Shalev-Arkushin, Vasileios Baltatzis, Connor Gillis et al.CHI 2025 · 12 citations
Builds on11
- TSPNet: Hierarchical Feature Learning via Temporal Semantic Pyramid for Sign Language TranslationDongxu Li, Chenchen Xu, Xin Yu, Kaihao Zhang et al.NeurIPS 2020 · 171 citations
- A Simple Multi-Modality Transfer Learning Baseline for Sign Language TranslationYutong Chen, Fangyun Wei, Xiao Sun, Zhirong Wu et al.CVPR 2022 · 137 citations
- Fingerspelling Recognition in the Wild With Iterative Visual AttentionBowen Shi, Aurora Martinez Del Rio, Jonathan Keane, Diane Brentari et al.ICCV 2019 · 76 citations
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 40 citations
- Aligning Subtitles in Sign Language VideosHannah Bull, Triantafyllos Afouras, Gül Varol, Samuel Albanie et al.ICCV 2021 · 39 citations
Related papers
- YouTube-SL-25: A Large-Scale, Open-Domain Multilingual Sign Language Parallel CorpusGarrett Tanzer, Biao ZhangICLR 2025
- Lost in Translation, Found in Context: Sign Language Translation with Contextual CuesYoungjoon Jang, Haran Raajesh, Liliane Momeni, Gül Varol et al.CVPR 2025
- Sign Language Video Retrieval with Free-Form Textual QueriesAmanda Cardoso Duarte, Samuel Albanie, Xavier Giró-i-Nieto, Gül VarolCVPR 2022 · 27 citations
- Uni-Sign: Toward Unified Sign Language Understanding at ScaleZecheng Li, Wengang Zhou, Weichao Zhao, Kepeng Wu et al.ICLR 2025
- Leveraging the Power of MLLMs for Gloss-Free Sign Language TranslationJungeun Kim, Hyeongwoo Jeon, Jongseong Bae, Ha Young KimICCV 2025 · 10 citations
