SynthVSR: Scaling Up Visual Speech RecognitionWith Synthetic Supervision
Xubo Liu, Egor Lakomkin, Konstantinos Vougioukas, Pingchuan Ma, Honglie Chen, Ruiming Xie, Morrie Doulaty, Niko Moritz, Jáchym Kolár, Stavros Petridis, Maja Pantic, Christian Fuegen
摘要
How are you?" "So cool!" * Work done during an internship at Meta AI. outperforming off-the-shelf approaches using thousands of hours of video. The WER is further reduced to 27.9% when using all 438 hours of labeled data from LRS3, which is on par with the state-of-the-art self-supervised AV-HuBERT method. Furthermore, when combined with large-scale pseudo-labeled audio-visual data SynthVSR yields a new state-of-the-art VSR WER of 16.9% using publicly available data only, surpassing the recent state-of-the-art approaches trained with 29 times more non-public machine-transcribed video data (90,000 hours). Finally, we perform extensive ablation studies to understand the effect of each component in our proposed method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Unified Speech Recognition: A Single Model for Auditory, Visual, and Audiovisual InputsAlexandros Haliassos, Rodrigo Mira, Honglie Chen, Zoe Landgraf 等NeurIPS 2024 · 被引用 22 次
- VALLR: Visual ASR Language Model for Lip ReadingMarshall Thomas, Edward Fish, Richard BowdenICCV 2025 · 被引用 6 次
- AudioVSR: Enhancing Video Speech Recognition with Audio DataXiaoda Yang, Xize Cheng, Jiaqi Duan, Hongshun Qiu 等EMNLP 2024 · 被引用 3 次
- Not Only Vision: Evolve Visual Speech Recognition via Peripheral InformationZhaoxin Yuan, Shuang Yang, Shiguang Shan, Xilin ChenICCV 2025 · 被引用 1 次
- Pay Attention to CTC: Fast and Robust Pseudo-Labelling for Unified Speech RecognitionAlexandros Haliassos, Rodrigo Mira, Stavros PetridisICLR 2026 · 被引用 1 次
它引用的顶会 Paper10
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- FastSpeech 2: Fast and High-Quality End-to-End Text to SpeechYi Ren, Chenxu Hu, Xu Tan, Tao Qin 等ICLR 2021 · 被引用 513 次
- Learning Audio-Visual Speech Representation by Masked Multimodal Cluster PredictionBowen Shi, Wei-Ning Hsu, Kushal Lakhotia, Abdelrahman MohamedICLR 2022 · 被引用 460 次
- Sub-word Level Lip Reading With Visual AttentionK. R. Prajwal, Triantafyllos Afouras, Andrew ZissermanCVPR 2022 · 被引用 104 次
- Spatio-Temporal Fusion Based Convolutional Sequence Learning for Lip ReadingXingxuan Zhang, Feng Cheng, Shilin WangICCV 2019 · 被引用 87 次
相关 Paper
- Jointly Learning Visual and Auditory Speech Representations from Raw DataAlexandros Haliassos, Pingchuan Ma, Rodrigo Mira, Stavros Petridis 等ICLR 2023 · 被引用 13 次
- u-HuBERT: Unified Mixed-Modal Speech Pretraining And Zero-Shot Transfer to Unlabeled ModalityWei-Ning Hsu, Bowen ShiNeurIPS 2022 · 被引用 58 次
- Lip2Vec: Efficient and Robust Visual Speech Recognition via Latent-to-Latent Visual to Audio Representation MappingYasser Abdelaziz Dahou Djilali, Sanath Narayan, Haithem Boussaid, Ebtesam Almazrouei 等ICCV 2023 · 被引用 17 次
- Towards Robust Speech Representation Learning for Thousands of LanguagesWilliam Chen, Wangyou Zhang, Yifan Peng, Xinjian Li 等EMNLP 2024 · 被引用 19 次
- Language-Guided Audio-Visual Source Separation via Trimodal ConsistencyReuben Tan, Arijit Ray, Andrea Burns, Bryan A. Plummer 等CVPR 2023
