Lip Reading for Low-resource Languages by Learning and Combining General Speech Knowledge and Language-specific Knowledge
Minsu Kim, Jeong Hun Yeo, Jeongsoo Choi, Yong Man Ro
Abstract
This paper proposes a novel lip reading framework, especially for low-resource languages, which has not been well addressed in the previous literature. Since low-resource languages do not have enough video-text paired data to train the model to have sufficient power to model lip movements and language, it is regarded as challenging to develop lip reading models for low-resource languages. In order to mitigate the challenge, we try to learn general speech knowledge, the ability to model lip movements, from a high-resource language through the prediction of speech units. It is known that different languages partially share common phonemes, thus general speech knowledge learned from one language can be extended to other languages. Then, we try to learn language-specific knowledge, the ability to model language, by proposing Language-specific Memory-augmented Decoder (LMDecoder). LMDecoder saves language-specific audio features into memory banks and can be trained on audio-text paired data which is more easily accessible than video-text paired data. Therefore, with LMDecoder, we can transform the input speech units into language-specific audio features and translate them into texts by utilizing the learned rich language knowledge. Finally, by combining general speech knowledge and language-specific knowledge, we can efficiently develop lip reading models even for low-resource languages. Through extensive experiments using five languages, English, Spanish, French, Italian, and Portuguese, the effectiveness of the proposed method is evaluated.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7c305f52-d95c-4ed1-b385-e95ea436b822Cited by top-tier papers4
- Personalized Lip Reading: Adapting to Your Unique Lip Movements with Vision and LanguageJeong Hun Yeo, Chae Won Kim, Hyunjun Kim, Hyeongseop Rha et al.AAAI 2025 · 7 citations
- Efficient Training for Multilingual Visual Speech Recognition: Pre-training with Discretized Visual Speech RepresentationMinsu Kim, Jeong Hun Yeo, Se Jin Park, Hyeongseop Rha et al.ACM MM 2024 · 4 citations
- Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech RepresentationsJeong Hun Yeo, Minsu Kim, Chae Won Kim, Stavros Petridis et al.ICCV 2025 · 3 citations
- AV2AV: Direct Audio-Visual Speech to Audio-Visual Speech Translation with Unified Audio-Visual Speech RepresentationJeongsoo Choi, Se Jin Park, Minsu Kim, Yong Man RoCVPR 2024
Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy et al.ICCV 2019 · 1,396 citations
- Unified Vision-Language Pre-Training for Image Captioning and VQALuowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu et al.AAAI 2020 · 1,047 citations
- vq-wav2vec: Self-Supervised Learning of Discrete Speech RepresentationsAlexei Baevski, Steffen Schneider, Michael AuliICLR 2020 · 730 citations
Related papers
- Distinguishing Homophenes Using Multi-Head Visual-Audio Memory for Lip ReadingMinsu Kim, Jeong Hun Yeo, Yong Man RoAAAI 2022 · 86 citations
- Multi-modality Associative Bridging through Memory: Speech Sound Recollected from Face VideoMinsu Kim, Joanna Hong, Se Jin Park, Yong Man RoICCV 2021 · 48 citations
- Hearing Lips: Improving Lip Reading by Distilling Speech RecognizersYa Zhao, Rui Xu, Xinchao Wang, Peng Hou et al.AAAI 2020 · 106 citations
- DualLip: A System for Joint Lip Reading and GenerationWeicong Chen, Xu Tan, Yingce Xia, Tao Qin et al.ACM MM 2020 · 25 citations
- CALLip: Lipreading using Contrastive and Attribute LearningYiyang Huang, Xuefeng Liang, Chaowei FangACM MM 2021 · 14 citations
