C2SLR: Consistency-enhanced Continuous Sign Language Recognition
Ronglai Zuo, Brian Mak
Abstract
The backbone of most deep-learning-based continuous sign language recognition (CSLR) models consists of a visual module, a sequential module, and an alignment module. However, such CSLR backbones are hard to be trained sufficiently with a single connectionist temporal classification loss. In this work, we propose two auxiliary constraints to enhance the CSLR backbones from the perspective of consistency. The first constraint aims to enhance the visual module, which easily suffers from the insufficient training problem. Specifically, since sign languages convey information mainly with signers' faces and hands, we insert a keypoint-guided spatial attention module into the visual module to enforce it to focus on informative regions, i.e., spatial attention consistency. Nevertheless, only enhancing the visual module may not fully exploit the power of the backbone. Motivated by that both the output features of the visual and sequential modules represent the same sentence, we further impose a sentence embedding consistency constraint between them to enhance the representation power of both the features. Experimental results over three representative backbones validate the effectiveness of the two constraints. More remarkably, with a transformer-based backbone, our model achieves state-of-the-art or competitive performance on three benchmarks, PHOENIX-2014, PHOENIX-2014-T, and CSL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e5f482cf-ab3b-449a-aa0e-c1b9b75655d0Cited by top-tier papers19
- Two-Stream Network for Sign Language Recognition and TranslationYutong Chen, Ronglai Zuo, Fangyun Wei, Yu Wu et al.NeurIPS 2022 · 288 citations
- Gloss-free Sign Language Translation: Improving from Visual-Language PretrainingBenjia Zhou, Zhigang Chen, Albert Clapés, Jun Wan et al.ICCV 2023 · 123 citations
- Self-Emphasizing Network for Continuous Sign Language RecognitionLianyu Hu, Liqing Gao, Zekang Liu, Wei FengAAAI 2023 · 91 citations
- CoSign: Exploring Co-occurrence Signals in Skeleton-based Continuous Sign Language RecognitionPeiqi Jiao, Yuecong Min, Yanan Li, Xiaotao Wang et al.ICCV 2023 · 52 citations
- Improving Continuous Sign Language Recognition with Cross-Lingual SignsFangyun Wei, Yutong ChenICCV 2023 · 46 citations
Builds on9
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- Spatial-Temporal Multi-Cue Network for Continuous Sign Language RecognitionHao Zhou, Wengang Zhou, Yun Zhou, Houqiang LiAAAI 2020 · 249 citations
- Mask-Guided Attention Network for Occluded Pedestrian DetectionYanwei Pang, Jin Xie, Muhammad Haris Khan, Rao Muhammad Anwer et al.ICCV 2019 · 216 citations
- Visual Alignment Constraint for Continuous Sign Language RecognitionYuecong Min, Aiming Hao, Xiujuan Chai, Xilin ChenICCV 2021 · 211 citations
- Self-Mutual Distillation Learning for Continuous Sign Language RecognitionAiming Hao, Yuecong Min, Xilin ChenICCV 2021 · 158 citations
Related papers
- Cross-Sentence Gloss Consistency for Continuous Sign Language RecognitionQi Rao, Ke Sun, Xiaohan Wang, Qi Wang et al.AAAI 2024 · 5 citations
- C2ST: Cross-modal Contextualized Sequence Transduction for Continuous Sign Language RecognitionHuaiwen Zhang, Zihang Guo, Yang Yang, Xin Liu et al.ICCV 2023 · 21 citations
- TCNet: Continuous Sign Language Recognition from Trajectories and Correlated RegionsHui Lu, Albert Ali Salah, Ronald PoppeAAAI 2024 · 20 citations
- Sign Language Transformers: Joint End-to-End Sign Language Recognition and TranslationNecati Cihan Camgöz, Oscar Koller, Simon Hadfield, Richard BowdenCVPR 2020
- Continuous Sign Language Recognition with Correlation NetworkLianyu Hu, Liqing Gao, Zekang Liu, Wei FengCVPR 2023
