Distilling Cross-Temporal Contexts for Continuous Sign Language Recognition
Leming Guo, Wanli Xue, Qing Guo, Bo Liu, Kaihua Zhang, Tiantian Yuan, Shengyong Chen
摘要
Continuous sign language recognition (CSLR) aims to recognize glosses in a sign language video. State-of-theart methods typically have two modules, a spatial perception module and a temporal aggregation module, which are jointly learned end-to-end. Existing results in [9, 20, 25, 36] have indicated that, as the frontal component of the overall model, the spatial perception module used for spatial feature extraction tends to be insufficiently trained. In this paper, we first conduct empirical studies and show that a shallow temporal aggregation module allows more thorough training of the spatial perception module. However, a shallow temporal aggregation module cannot well capture both local and global temporal context information in sign language. To address this dilemma, we propose a crosstemporal context aggregation (CTCA) model. Specifically, we build a dual-path network that contains two branches for perceptions of local temporal context and global temporal context. We further design a cross-context knowledge distillation learning objective to aggregate the two types of context and the linguistic prior. The knowledge distillation enables the resultant one-branch temporal aggregation module to perceive local-global temporal and semantic context. This shallow temporal perception module structure facilitates spatial perception module learning. Extensive experiments on challenging CSLR benchmarks demonstrate that our method outperforms all state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- TCNet: Continuous Sign Language Recognition from Trajectories and Correlated RegionsHui Lu, Albert Ali Salah, Ronald PoppeAAAI 2024 · 被引用 20 次
- Towards Online Continuous Sign Language Recognition and TranslationRonglai Zuo, Fangyun Wei, Brian MakEMNLP 2024 · 被引用 14 次
- OLMD: Orientation-aware Long-term Motion Decoupling for Continuous Sign Language RecognitionYiheng Yu, Sheng Liu, Yuan Feng, Min Xu 等AAAI 2025 · 被引用 5 次
- SEDS: Semantically Enhanced Dual-Stream Encoder for Sign Language RetrievalLongtao Jiang, Min Wang, Zecheng Li, Yao Fang 等ACM MM 2024 · 被引用 2 次
- HyperSign: Hierarchical Hypergraph-based Co-occurrence Modeling for Sign Language Recognition and TranslationQianren Guo, Yuehang Wang, Yongji Zhang, Qi Chu 等AAAI 2026
它引用的顶会 Paper16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Channel-wise Knowledge Distillation for Dense Prediction*Changyong Shu, Yifan Liu, Jianfei Gao, Zheng Yan 等ICCV 2021 · 被引用 432 次
- Deep Multimodal Fusion by Channel ExchangingYikai Wang, Wenbing Huang, Fuchun Sun, Tingyang Xu 等NeurIPS 2020 · 被引用 321 次
- Spatial-Temporal Multi-Cue Network for Continuous Sign Language RecognitionHao Zhou, Wengang Zhou, Yun Zhou, Houqiang LiAAAI 2020 · 被引用 249 次
- Visual Alignment Constraint for Continuous Sign Language RecognitionYuecong Min, Aiming Hao, Xiujuan Chai, Xilin ChenICCV 2021 · 被引用 211 次
相关 Paper
- Self-Mutual Distillation Learning for Continuous Sign Language RecognitionAiming Hao, Yuecong Min, Xilin ChenICCV 2021 · 被引用 158 次
- C2ST: Cross-modal Contextualized Sequence Transduction for Continuous Sign Language RecognitionHuaiwen Zhang, Zihang Guo, Yang Yang, Xin Liu 等ICCV 2023 · 被引用 21 次
- MixSignGraph: A Sign Sequence is Worth Mixed Graphs of NodesShiwei Gan, Yafeng Yin, Zhiwei Jiang, Lei Xie 等NeurIPS 2025 · 被引用 11 次
- Conditional Variational Autoencoder for Sign Language Translation with Cross-Modal AlignmentRui Zhao, Liang Zhang, Biao Fu, Cong Hu 等AAAI 2024 · 被引用 36 次
- Sign Language Transformers: Joint End-to-End Sign Language Recognition and TranslationNecati Cihan Camgöz, Oscar Koller, Simon Hadfield, Richard BowdenCVPR 2020
