Spatio-Temporal Fusion Based Convolutional Sequence Learning for Lip Reading
Xingxuan Zhang, Feng Cheng, Shilin Wang
摘要
Current state-of-the-art approaches for lip reading are based on sequence-to-sequence architectures that are designed for natural machine translation and audio speech recognition. Hence, these methods do not fully exploit the characteristics of the lip dynamics, causing two main drawbacks. First, the short-range temporal dependencies, which are critical to the mapping from lip images to visemes, receives no extra attention. Second, local spatial information is discarded in the existing sequence models due to the use of global average pooling (GAP). To well solve these drawbacks, we propose a Temporal Focal block to sufficiently describe short-range dependencies and a Spatio-Temporal Fusion Module (STFM) to maintain the local spatial information and to reduce the feature dimensions as well. From the experiment results, it is demonstrated that our method achieves comparable performance with the state-of-the-art approach using much less training data and much lighter Convolutional Feature Extractor. The training time is reduced by 12 days due to the convolutional structure and the local self-attention mechanism.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Learning Audio-Visual Speech Representation by Masked Multimodal Cluster PredictionBowen Shi, Wei-Ning Hsu, Kushal Lakhotia, Abdelrahman MohamedICLR 2022 · 被引用 460 次
- Sub-word Level Lip Reading With Visual AttentionK. R. Prajwal, Triantafyllos Afouras, Andrew ZissermanCVPR 2022 · 被引用 104 次
- Distinguishing Homophenes Using Multi-Head Visual-Audio Memory for Lip ReadingMinsu Kim, Jeong Hun Yeo, Yong Man RoAAAI 2022 · 被引用 86 次
- Contrastive Learning of Global and Local Video RepresentationsShuang Ma, Zhaoyang Zeng, Daniel McDuff, Yale SongNeurIPS 2021 · 被引用 45 次
- Leveraging Unimodal Self-Supervised Learning for Multimodal Audio-Visual Speech RecognitionXichen Pan, Peiyu Chen, Yichen Gong, Helong Zhou 等ACL 2022 · 被引用 43 次
相关 Paper
- Multi-Perspective LSTM for Joint Visual Representation LearningAlireza Sepas-Moghaddam, Fernando Pereira, Paulo Lobato Correia, Ali EtemadCVPR 2021
- Hearing Lips: Improving Lip Reading by Distilling Speech RecognizersYa Zhao, Rui Xu, Xinchao Wang, Peng Hou 等AAAI 2020 · 被引用 106 次
- EventLip: Enhancing Event-Based Lip Reading via Frequency-Aware Spatiotemporal Hypergraph ModelingXueyi Zhang, Jialu Sun, Chengwei Zhang, Xianghu Yue 等ACM MM 2025
- Gait Recognition via Effective Global-Local Feature Representation and Local Temporal AggregationBeibei Lin, Shunli Zhang, Xin YuICCV 2021 · 被引用 325 次
- Discriminative Multi-Modality Speech RecognitionBo Xu, Cheng Lu, Yandong Guo, Jacob WangCVPR 2020
