Cuing Without Sharing: A Federated Cued Speech Recognition Framework via Mutual Knowledge Distillation
Yuxuan Zhang, Lei Liu, Li Liu
摘要
Cued Speech (CS) is a visual coding tool to encode spoken languages at the phonetic level, which combines lip-reading and hand gestures to effectively assist communication among people with hearing impairments. The Automatic CS Recognition (ACSR) task aims to recognize CS videos into linguistic texts, which involves both lips and hands as two distinct modalities conveying complementary information. However, the traditional centralized training approach poses potential privacy risks due to the use of facial and gesture videos in CS data. To address this issue, we propose a new Federated Cued Speech Recognition (FedCSR) framework to train an ACSR model over the decentralized CS data without sharing private information. In particular, a mutual knowledge distillation method is proposed to maintain cross-modal semantic consistency of the Non-IID CS data, which ensures learning a unified feature space for both linguistic and visual information. On the server side, a globally shared linguistic model is trained to capture the long-term dependencies in the text sentences, which is aligned with the visual information from the local clients via visual-tolinguistic distillation. On the client side, the visual model of each client is trained with its own local data, assisted by linguistic-tovisual distillation treating the linguistic model as the teacher. To the best of our knowledge, this is the first approach to consider the federated ACSR task for privacy protection. Experimental results on the Chinese CS dataset with multiple cuers 1 demonstrate that our approach outperforms both mainstream federated learning baselines and existing centralized state-of-the-art ACSR methods, achieving 9.7% performance improvement for character error rate (CER) and 15.0% for word error rate (WER). Code is available at https://github.com/YuxuanZHANG0713/FedCSR.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Global Balanced Experts for Federated Long-Tailed LearningYaopei Zeng, Lei Liu, Li Liu, Li Shen 等ICCV 2023 · 被引用 10 次
- PFedCS: A Personalized Federated Learning Method for Enhancing Collaboration among Similar ClassifiersSiyuan Wu, Yongzhe Jia, Bowen Liu, Haolong Xiang 等AAAI 2025 · 被引用 6 次
- UniCUE: Unified Recognition and Generation Framework for Chinese Cued Speech Video-to-Speech GenerationJinting Wang, Shan Yang, Chenxing Li, Dong Yu 等AAAI 2026
它引用的顶会 Paper7
- SCAFFOLD: Stochastic Controlled Averaging for Federated LearningSai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J. Reddi 等ICML 2020 · 被引用 3,875 次
- FedBN: Federated Learning on Non-IID Features via Local Batch NormalizationXiaoxiao Li, Meirui Jiang, Xiaofei Zhang, Michael Kamp 等ICLR 2021 · 被引用 1,166 次
- HarmoFL: Harmonizing Local and Global Drifts in Federated Learning on Heterogeneous Medical ImagesMeirui Jiang, Zirui Wang, Qi DouAAAI 2022 · 被引用 187 次
- Performance Optimization of Federated Person Re-identification via Benchmark AnalysisWeiming Zhuang, Yonggang Wen, Xuesen Zhang, Xin Gan 等ACM MM 2020 · 被引用 94 次
- Joint Optimization in Edge-Cloud Continuum for Federated Unsupervised Person Re-identificationWeiming Zhuang, Yonggang Wen, Shuai ZhangACM MM 2021 · 被引用 43 次
相关 Paper
- Cued-Agent: A Collaborative Multi-Agent System for Automatic Cued Speech RecognitionGuanjie Huang, Danny H. K. Tsang, Shan Yang, Guangzhi Lei 等ACM MM 2025
- Hearing Lips: Improving Lip Reading by Distilling Speech RecognizersYa Zhao, Rui Xu, Xinchao Wang, Peng Hou 等AAAI 2020 · 被引用 106 次
- Multi-Task Corrupted Prediction for Learning Robust Audio-Visual Speech RepresentationSungnyun Kim, Sungwoo Cho, Sangmin Bae, Kangwook Jang 等ICLR 2025
- Cueing Without Gapping: Cuer-Independent Cued Speech Recognition Powered by Cross-Cuer Invariant ModelingFengji Ma, Chenxing Li, Li LiuAAAI 2026
- Learning From the Master: Distilling Cross-Modal Advanced Knowledge for Lip ReadingSucheng Ren, Yong Du, Jianming Lv, Guoqiang Han 等CVPR 2021
