Self-Mutual Distillation Learning for Continuous Sign Language Recognition
Aiming Hao, Yuecong Min, Xilin Chen
摘要
In recent years, deep learning moves video-based Continuous Sign Language Recognition (CSLR) significantly forward. Currently, a typical network combination for CSLR includes a visual module, which focuses on spatial and short-temporal information, followed by a contextual module, which focuses on long-temporal information, and the Connectionist Temporal Classification (CTC) loss is adopted to train the network. However, due to the limitation of chain rules in back-propagation, the visual module is hard to adjust for seeking optimized visual features. As a result, it enforces that the contextual module focuses on contextual information optimization only rather than balancing efficient visual and contextual information. In this paper, we propose a Self-Mutual Knowledge Distillation (SMKD) method, which enforces the visual and contextual modules to focus on short-term and long-term information and enhances the discriminative power of both modules simultaneously. Specifically, the visual and contextual modules share the weights of their corresponding classifiers, and train with CTC loss simultaneously. Moreover, the spike phenomenon widely exists with CTC loss. Although it can help us choose a few of the key frames of a gloss, it does drop other frames in a gloss and makes the visual feature saturation in the early stage. A gloss segmentation is developed to relieve the spike phenomenon and decrease saturation in the visual module. We conduct experiments on two CSLR bench-marks: PHOENIX14 and PHOENIX14-T. Experimental results demonstrate the effectiveness of the SMKD.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- Two-Stream Network for Sign Language Recognition and TranslationYutong Chen, Ronglai Zuo, Fangyun Wei, Yu Wu 等NeurIPS 2022 · 被引用 288 次
- Gloss-free Sign Language Translation: Improving from Visual-Language PretrainingBenjia Zhou, Zhigang Chen, Albert Clapés, Jun Wan 等ICCV 2023 · 被引用 123 次
- C2SLR: Consistency-enhanced Continuous Sign Language RecognitionRonglai Zuo, Brian MakCVPR 2022 · 被引用 118 次
- Self-Emphasizing Network for Continuous Sign Language RecognitionLianyu Hu, Liqing Gao, Zekang Liu, Wei FengAAAI 2023 · 被引用 91 次
- CoSign: Exploring Co-occurrence Signals in Skeleton-based Continuous Sign Language RecognitionPeiqi Jiao, Yuecong Min, Yanan Li, Xiaotao Wang 等ICCV 2023 · 被引用 52 次
它引用的顶会 Paper7
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen 等ICCV 2019 · 被引用 1,069 次
- Spatial-Temporal Multi-Cue Network for Continuous Sign Language RecognitionHao Zhou, Wengang Zhou, Yun Zhou, Houqiang LiAAAI 2020 · 被引用 249 次
- Visual Alignment Constraint for Continuous Sign Language RecognitionYuecong Min, Aiming Hao, Xiujuan Chai, Xilin ChenICCV 2021 · 被引用 211 次
- Boosting Continuous Sign Language Recognition via Cross Modality AugmentationJunfu Pu, Wengang Zhou, Hezhen Hu, Houqiang LiACM MM 2020 · 被引用 121 次
- Sign Language Transformers: Joint End-to-End Sign Language Recognition and TranslationNecati Cihan Camgöz, Oscar Koller, Simon Hadfield, Richard BowdenCVPR 2020
相关 Paper
- C2ST: Cross-modal Contextualized Sequence Transduction for Continuous Sign Language RecognitionHuaiwen Zhang, Zihang Guo, Yang Yang, Xin Liu 等ICCV 2023 · 被引用 21 次
- Distilling Cross-Temporal Contexts for Continuous Sign Language RecognitionLeming Guo, Wanli Xue, Qing Guo, Bo Liu 等CVPR 2023
- TCNet: Continuous Sign Language Recognition from Trajectories and Correlated RegionsHui Lu, Albert Ali Salah, Ronald PoppeAAAI 2024 · 被引用 20 次
- Cross-Sentence Gloss Consistency for Continuous Sign Language RecognitionQi Rao, Ke Sun, Xiaohan Wang, Qi Wang 等AAAI 2024 · 被引用 5 次
- Conditional Variational Autoencoder for Sign Language Translation with Cross-Modal AlignmentRui Zhao, Liang Zhang, Biao Fu, Cong Hu 等AAAI 2024 · 被引用 36 次
