CALLip: Lipreading using Contrastive and Attribute Learning
Yiyang Huang, Xuefeng Liang, Chaowei Fang
摘要
Lipreading, aiming at interpreting speech by watching the lip movements of the speaker, has great significance in human communication and speech understanding. Despite having reached a feasible performance, lipreading still faces two crucial challenges: 1) the considerable lip movement variations cross different persons when they utter the same words; 2) the similar lip movements of people when they utter some confused phonemes. To tackle these two problems, we propose a novel lipreading framework, CALLip, which employs attribute learning and contrastive learning. The attribute learning extracts the speaker identity-aware features through a speaker recognition branch, which are able to normalize the lip shapes to eliminate cross-speaker variations. Considering that audio signals are intrinsically more distinguishable than visual signals, the contrastive learning is devised between visual and audio signals to enhance the discrimination of visual features and alleviate the viseme confusion problem. Experimental results show that CALLip does learn better features of lip movements. The comparisons on both English and Chinese benchmark datasets, GRID and CMLR, demonstrate that CALLip outperforms six state-of-the-art lipreading methods without using any additional data.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- LipLearner: Customizable Silent Speech Interactions on Mobile DevicesZixiong Su, Shitao Fang, Jun RekimotoCHI 2023 · 被引用 37 次
- LinProVSR: Linguistics-Knowledge Guided Progressive Disambiguation Network for Visual Speech RecognitionFeng Xue, Baochao Zhu, Wei Jia, Shujie Li 等AAAI 2026
相关 Paper
- Hearing Lips: Improving Lip Reading by Distilling Speech RecognizersYa Zhao, Rui Xu, Xinchao Wang, Peng Hou 等AAAI 2020 · 被引用 106 次
- Cross-modal Self-Supervised Learning for Lip Reading: When Contrastive Learning meets Adversarial TrainingChangchong Sheng, Matti Pietikäinen, Qi Tian, Li LiuACM MM 2021 · 被引用 11 次
- Distinguishing Homophenes Using Multi-Head Visual-Audio Memory for Lip ReadingMinsu Kim, Jeong Hun Yeo, Yong Man RoAAAI 2022 · 被引用 86 次
- Seeing What You Said: Talking Face Generation Guided by a Lip Reading ExpertJiadong Wang, Xinyuan Qian, Malu Zhang, Robby T. Tan 等CVPR 2023
- SyncTalklip: Highly Synchronized Lip-Readable Speaker Generation with Multi-Task LearningXiaoda Yang, Xize Cheng, Dongjie Fu, Minghui Fang 等ACM MM 2024 · 被引用 4 次
