Distilling Cross-Temporal Contexts for Continuous Sign Language Recognition
Leming Guo, Wanli Xue, Qing Guo, Bo Liu, Kaihua Zhang, Tiantian Yuan, Shengyong Chen
Abstract
Continuous sign language recognition (CSLR) aims to recognize glosses in a sign language video. State-of-theart methods typically have two modules, a spatial perception module and a temporal aggregation module, which are jointly learned end-to-end. Existing results in [9, 20, 25, 36] have indicated that, as the frontal component of the overall model, the spatial perception module used for spatial feature extraction tends to be insufficiently trained. In this paper, we first conduct empirical studies and show that a shallow temporal aggregation module allows more thorough training of the spatial perception module. However, a shallow temporal aggregation module cannot well capture both local and global temporal context information in sign language. To address this dilemma, we propose a crosstemporal context aggregation (CTCA) model. Specifically, we build a dual-path network that contains two branches for perceptions of local temporal context and global temporal context. We further design a cross-context knowledge distillation learning objective to aggregate the two types of context and the linguistic prior. The knowledge distillation enables the resultant one-branch temporal aggregation module to perceive local-global temporal and semantic context. This shallow temporal perception module structure facilitates spatial perception module learning. Extensive experiments on challenging CSLR benchmarks demonstrate that our method outperforms all state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b960e878-ba06-4497-8a3a-69d5a28607f8Cited by top-tier papers6
- TCNet: Continuous Sign Language Recognition from Trajectories and Correlated RegionsHui Lu, Albert Ali Salah, Ronald PoppeAAAI 2024 · 20 citations
- Towards Online Continuous Sign Language Recognition and TranslationRonglai Zuo, Fangyun Wei, Brian MakEMNLP 2024 · 14 citations
- OLMD: Orientation-aware Long-term Motion Decoupling for Continuous Sign Language RecognitionYiheng Yu, Sheng Liu, Yuan Feng, Min Xu et al.AAAI 2025 · 5 citations
- SEDS: Semantically Enhanced Dual-Stream Encoder for Sign Language RetrievalLongtao Jiang, Min Wang, Zecheng Li, Yao Fang et al.ACM MM 2024 · 2 citations
- HyperSign: Hierarchical Hypergraph-based Co-occurrence Modeling for Sign Language Recognition and TranslationQianren Guo, Yuehang Wang, Yongji Zhang, Qi Chu et al.AAAI 2026
Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Channel-wise Knowledge Distillation for Dense Prediction*Changyong Shu, Yifan Liu, Jianfei Gao, Zheng Yan et al.ICCV 2021 · 432 citations
- Deep Multimodal Fusion by Channel ExchangingYikai Wang, Wenbing Huang, Fuchun Sun, Tingyang Xu et al.NeurIPS 2020 · 321 citations
- Spatial-Temporal Multi-Cue Network for Continuous Sign Language RecognitionHao Zhou, Wengang Zhou, Yun Zhou, Houqiang LiAAAI 2020 · 249 citations
- Visual Alignment Constraint for Continuous Sign Language RecognitionYuecong Min, Aiming Hao, Xiujuan Chai, Xilin ChenICCV 2021 · 211 citations
Related papers
- Self-Mutual Distillation Learning for Continuous Sign Language RecognitionAiming Hao, Yuecong Min, Xilin ChenICCV 2021 · 158 citations
- C2ST: Cross-modal Contextualized Sequence Transduction for Continuous Sign Language RecognitionHuaiwen Zhang, Zihang Guo, Yang Yang, Xin Liu et al.ICCV 2023 · 21 citations
- MixSignGraph: A Sign Sequence is Worth Mixed Graphs of NodesShiwei Gan, Yafeng Yin, Zhiwei Jiang, Lei Xie et al.NeurIPS 2025 · 11 citations
- Conditional Variational Autoencoder for Sign Language Translation with Cross-Modal AlignmentRui Zhao, Liang Zhang, Biao Fu, Cong Hu et al.AAAI 2024 · 36 citations
- Sign Language Transformers: Joint End-to-End Sign Language Recognition and TranslationNecati Cihan Camgöz, Oscar Koller, Simon Hadfield, Richard BowdenCVPR 2020
