Cuing Without Sharing: A Federated Cued Speech Recognition Framework via Mutual Knowledge Distillation
Yuxuan Zhang, Lei Liu, Li Liu
Abstract
Cued Speech (CS) is a visual coding tool to encode spoken languages at the phonetic level, which combines lip-reading and hand gestures to effectively assist communication among people with hearing impairments. The Automatic CS Recognition (ACSR) task aims to recognize CS videos into linguistic texts, which involves both lips and hands as two distinct modalities conveying complementary information. However, the traditional centralized training approach poses potential privacy risks due to the use of facial and gesture videos in CS data. To address this issue, we propose a new Federated Cued Speech Recognition (FedCSR) framework to train an ACSR model over the decentralized CS data without sharing private information. In particular, a mutual knowledge distillation method is proposed to maintain cross-modal semantic consistency of the Non-IID CS data, which ensures learning a unified feature space for both linguistic and visual information. On the server side, a globally shared linguistic model is trained to capture the long-term dependencies in the text sentences, which is aligned with the visual information from the local clients via visual-tolinguistic distillation. On the client side, the visual model of each client is trained with its own local data, assisted by linguistic-tovisual distillation treating the linguistic model as the teacher. To the best of our knowledge, this is the first approach to consider the federated ACSR task for privacy protection. Experimental results on the Chinese CS dataset with multiple cuers 1 demonstrate that our approach outperforms both mainstream federated learning baselines and existing centralized state-of-the-art ACSR methods, achieving 9.7% performance improvement for character error rate (CER) and 15.0% for word error rate (WER). Code is available at https://github.com/YuxuanZHANG0713/FedCSR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a6d8b7eb-db35-4d1f-9c09-54528914a1f2Cited by top-tier papers3
- Global Balanced Experts for Federated Long-Tailed LearningYaopei Zeng, Lei Liu, Li Liu, Li Shen et al.ICCV 2023 · 10 citations
- PFedCS: A Personalized Federated Learning Method for Enhancing Collaboration among Similar ClassifiersSiyuan Wu, Yongzhe Jia, Bowen Liu, Haolong Xiang et al.AAAI 2025 · 6 citations
- UniCUE: Unified Recognition and Generation Framework for Chinese Cued Speech Video-to-Speech GenerationJinting Wang, Shan Yang, Chenxing Li, Dong Yu et al.AAAI 2026
Builds on7
- SCAFFOLD: Stochastic Controlled Averaging for Federated LearningSai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J. Reddi et al.ICML 2020 · 3,875 citations
- FedBN: Federated Learning on Non-IID Features via Local Batch NormalizationXiaoxiao Li, Meirui Jiang, Xiaofei Zhang, Michael Kamp et al.ICLR 2021 · 1,166 citations
- HarmoFL: Harmonizing Local and Global Drifts in Federated Learning on Heterogeneous Medical ImagesMeirui Jiang, Zirui Wang, Qi DouAAAI 2022 · 187 citations
- Performance Optimization of Federated Person Re-identification via Benchmark AnalysisWeiming Zhuang, Yonggang Wen, Xuesen Zhang, Xin Gan et al.ACM MM 2020 · 94 citations
- Joint Optimization in Edge-Cloud Continuum for Federated Unsupervised Person Re-identificationWeiming Zhuang, Yonggang Wen, Shuai ZhangACM MM 2021 · 43 citations
Related papers
- Cued-Agent: A Collaborative Multi-Agent System for Automatic Cued Speech RecognitionGuanjie Huang, Danny H. K. Tsang, Shan Yang, Guangzhi Lei et al.ACM MM 2025
- Hearing Lips: Improving Lip Reading by Distilling Speech RecognizersYa Zhao, Rui Xu, Xinchao Wang, Peng Hou et al.AAAI 2020 · 106 citations
- Multi-Task Corrupted Prediction for Learning Robust Audio-Visual Speech RepresentationSungnyun Kim, Sungwoo Cho, Sangmin Bae, Kangwook Jang et al.ICLR 2025
- Cueing Without Gapping: Cuer-Independent Cued Speech Recognition Powered by Cross-Cuer Invariant ModelingFengji Ma, Chenxing Li, Li LiuAAAI 2026
- Learning From the Master: Distilling Cross-Modal Advanced Knowledge for Lip ReadingSucheng Ren, Yong Du, Jianming Lv, Guoqiang Han et al.CVPR 2021
