Two-Stream Network for Sign Language Recognition and Translation
Yutong Chen, Ronglai Zuo, Fangyun Wei, Yu Wu, Shujie Liu, Brian Mak
Abstract
Sign languages are visual languages using manual articulations and non-manual elements to convey information. For sign language recognition and translation, the majority of existing approaches directly encode RGB videos into hidden representations. RGB videos, however, are raw signals with substantial visual redundancy, leading the encoder to overlook the key information for sign language understanding. To mitigate this problem and better incorporate domain knowledge, such as handshape and body movement, we introduce a dual visual encoder containing two separate streams to model both the raw videos and the keypoint sequences generated by an off-the-shelf keypoint estimator. To make the two streams interact with each other, we explore a variety of techniques, including bidirectional lateral connection, sign pyramid network with auxiliary supervision, and frame-level selfdistillation. The resulting model is called TwoStream-SLR, which is competent for sign language recognition (SLR). TwoStream-SLR is extended to a sign language translation (SLT) model, TwoStream-SLT, by simply attaching an extra translation network. Experimentally, our TwoStream-SLR and TwoStream-SLT achieve stateof-the-art performance on SLR and SLT tasks across a series of datasets including Phoenix-2014, Phoenix-2014T, and CSL-Daily. Code and models are available at: https://github.com/FangyunWei/SLRT . * Equal contribution. † Corresponding author. 1 Glosses are the word-for-word transcription of sign language where each gloss is a unique label for a sign. 36th Conference on Neural Information Processing Systems (NeurIPS 2022).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9f01c2cb-5cc5-41fd-90be-036e3484e4c2Cited by top-tier papers34
- Gloss-free Sign Language Translation: Improving from Visual-Language PretrainingBenjia Zhou, Zhigang Chen, Albert Clapés, Jun Wan et al.ICCV 2023 · 123 citations
- Sign2GPT: Leveraging Large Language Models for Gloss-Free Sign Language TranslationRyan Wong, Necati Cihan Camgöz, Richard BowdenICLR 2024 · 58 citations
- CoSign: Exploring Co-occurrence Signals in Skeleton-based Continuous Sign Language RecognitionPeiqi Jiao, Yuecong Min, Yanan Li, Xiaotao Wang et al.ICCV 2023 · 52 citations
- Improving Continuous Sign Language Recognition with Cross-Lingual SignsFangyun Wei, Yutong ChenICCV 2023 · 46 citations
- Sign Language Translation with Iterative PrototypeHuijie Yao, Wengang Zhou, Hao Feng, Hezhen Hu et al.ICCV 2023 · 31 citations
Builds on14
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei et al.CVPR 2022 · 1,847 citations
- Revisiting Skeleton-based Action RecognitionHaodong Duan, Yue Zhao, Kai Chen, Dahua Lin et al.CVPR 2022 · 752 citations
- Visual Alignment Constraint for Continuous Sign Language RecognitionYuecong Min, Aiming Hao, Xiujuan Chai, Xilin ChenICCV 2021 · 211 citations
Related papers
- Natural Language-Assisted Sign Language RecognitionRonglai Zuo, Fangyun Wei, Brian MakCVPR 2023
- Conditional Variational Autoencoder for Sign Language Translation with Cross-Modal AlignmentRui Zhao, Liang Zhang, Biao Fu, Cong Hu et al.AAAI 2024 · 36 citations
- BoostSLT: Boosting Sign Language Translation via a Plug-and-Play Diffusion-Based Semantic EnhancerChangzhou Han, Wanlun Ma, Xi Tang, Kun Hu et al.CVPR 2026
- TSPNet: Hierarchical Feature Learning via Temporal Semantic Pyramid for Sign Language TranslationDongxu Li, Chenchen Xu, Xin Yu, Kaihao Zhang et al.NeurIPS 2020 · 171 citations
- SEDS: Semantically Enhanced Dual-Stream Encoder for Sign Language RetrievalLongtao Jiang, Min Wang, Zecheng Li, Yao Fang et al.ACM MM 2024 · 2 citations
