SHuBERT: Self-Supervised Sign Language Representation Learning via Multi-Stream Cluster Prediction
Shester Gueuwou, Xiaodan Du, Greg Shakhnarovich, Karen Livescu, Alexander H. Liu
Abstract
Sign language processing has traditionally relied on task-specific models, limiting the potential for transfer learning across tasks. Pretraining methods for sign language have typically focused on either supervised pre-training, which cannot take advantage of unlabeled data, or context-independent (frame or video segment) representations, which ignore the effects of relationships across time in sign language. We introduce SHuBERT (Sign Hidden-Unit BERT), a self-supervised contextual representation model learned from approximately 1,000 hours of American Sign Language video. SHu-BERT adapts masked token prediction objectives to multi-stream visual sign language input, learning to predict multiple targets corresponding to clustered hand, face, and body pose streams. SHuBERT achieves state-of-theart performance across multiple tasks including sign language translation, isolated sign language recognition, and fingerspelling detection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 857caeb3-ca11-42b4-9588-054e15eafee7Cited by top-tier papers2
- Bridging Sign and Spoken Languages: Pseudo Gloss Generation for Sign Language TranslationJianyuan Guo, Peike Li, Trevor CohnNeurIPS 2025 · 17 citations
- Learning Effective Sign Features without Text for Gloss-free Sign Language TranslationShiwei Gan, Xiao Liu, Yafeng Yin, Nan Liu et al.CVPR 2026 · 2 citations
Builds on17
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- SignBERT: Pre-Training of Hand-Model-Aware Representation for Sign Language RecognitionHezhen Hu, Weichao Zhao, Wengang Zhou, Yuechen Wang et al.ICCV 2021 · 125 citations
- Fingerspelling Recognition in the Wild With Iterative Visual AttentionBowen Shi, Aurora Martinez Del Rio, Jonathan Keane, Diane Brentari et al.ICCV 2019 · 76 citations
- Sign2GPT: Leveraging Large Language Models for Gloss-Free Sign Language TranslationRyan Wong, Necati Cihan Camgöz, Richard BowdenICLR 2024 · 58 citations
Related papers
- BEST: BERT Pre-training for Sign Language Recognition with Coupling TokenizationWeichao Zhao, Hezhen Hu, Wengang Zhou, Jiaxin Shi et al.AAAI 2023 · 70 citations
- Learning Audio-Visual Speech Representation by Masked Multimodal Cluster PredictionBowen Shi, Wei-Ning Hsu, Kushal Lakhotia, Abdelrahman MohamedICLR 2022 · 460 citations
- Scaling Sign Language TranslationBiao Zhang, Garrett Tanzer, Orhan FiratNeurIPS 2024 · 21 citations
- Aligning Subtitles in Sign Language VideosHannah Bull, Triantafyllos Afouras, Gül Varol, Samuel Albanie et al.ICCV 2021 · 39 citations
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy et al.ICCV 2019 · 1,396 citations
