Multi-grained Spatio-Temporal Features Perceived Network for Event-based Lip-Reading
Ganchao Tan, Yang Wang, Han Han, Yang Cao, Feng Wu, Zhengjun Zha
摘要
Automatic lip-reading (ALR) aims to recognize words using visual information from the speaker's lip movements. In this work, we introduce a novel type of sensing device, event cameras, for the task of ALR. Event cameras have both technical and application advantages over conventional cameras for the ALR task because they have higher temporal resolution, less redundant visual information, and lower power consumption. To recognize words from the event data, we propose a novel Multi-grained Spatio-Temporal Features Perceived Network (MSTP) to perceive fine-grained spatio-temporal features from microsecond time-resolved event data. Specifically, a multi-branch network architecture is designed, in which different grained spatio-temporal features are learned by operating at different frame rates. The branch operating on the low frame rate can perceive spatial complete but temporal coarse features. While the branch operating on the high frame rate can perceive spatial coarse but temporal refinement features. And a message flow module is devised to integrate the features from different branches, leading to perceiving more discriminative spatio-temporal features. In addition, we present the first event-based lip-reading dataset (DVS-Lip) captured by the event camera. Experimental results demonstrated the superiority of the proposed model compared to the state-of-the-art event-based action recognition models and video-based lip-reading models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Dilated convolution with learnable spacingsIsmail Khalfaoui Hassani, Thomas Pellegrini, Timothée MasquelierICLR 2023 · 被引用 18 次
- MTGA: Multi-View Temporal Granularity Aligned Aggregation for Event-Based Lip-ReadingWenhao Zhang, Jun Wang, Yong Luo, Lei Yu 等AAAI 2025 · 被引用 8 次
- Fractional-Order Spiking Neural NetworkChengjie Ge, Yufeng Peng, Zihao Li, Qiyu Kang 等ICLR 2026 · 被引用 5 次
- Multiplication-Free Parallelizable Spiking Neurons with Efficient Spatio-Temporal DynamicsPeng Xue, Wei Fang, Zhengyu Ma, Zihan Huang 等NeurIPS 2025 · 被引用 5 次
- Event2Vec: Processing neuromorphic events directly by representations in vector spaceWei Fang, Priyadarshini PandaICML 2026 · 被引用 4 次
它引用的顶会 Paper8
- ViTAE: Vision Transformer Advanced by Exploring Intrinsic Inductive BiasYufei Xu, Qiming Zhang, Jing Zhang, Dacheng TaoNeurIPS 2021 · 被引用 429 次
- End-to-End Learning of Representations for Asynchronous Event-Based DataDaniel Gehrig, Antonio Loquercio, Konstantinos G. Derpanis, Davide ScaramuzzaICCV 2019 · 被引用 427 次
- DualLip: A System for Joint Lip Reading and GenerationWeicong Chen, Xu Tan, Yingce Xia, Tao Qin 等ACM MM 2020 · 被引用 25 次
- SimulLR: Simultaneous Lip Reading Transducer with Attention-Guided Adaptive MemoryZhijie Lin, Zhou Zhao, Haoyuan Li, Jinglin Liu 等ACM MM 2021 · 被引用 13 次
- Time Lens: Event-Based Video Frame InterpolationStepan Tulyakov, Daniel Gehrig, Stamatios Georgoulis, Julius Erbach 等CVPR 2021
相关 Paper
- EventLip: Enhancing Event-Based Lip Reading via Frequency-Aware Spatiotemporal Hypergraph ModelingXueyi Zhang, Jialu Sun, Chengwei Zhang, Xianghu Yue 等ACM MM 2025
- Discriminative Multi-Modality Speech RecognitionBo Xu, Cheng Lu, Yandong Guo, Jacob WangCVPR 2020
- Hearing Lips: Improving Lip Reading by Distilling Speech RecognizersYa Zhao, Rui Xu, Xinchao Wang, Peng Hou 等AAAI 2020 · 被引用 106 次
- Learning an Event Sequence Embedding for Dense Event-Based Deep StereoStepan Tulyakov, François Fleuret, Martin Kiefel, Peter V. Gehler 等ICCV 2019 · 被引用 122 次
- LLaFEA: Frame-Event Complementary Fusion for Fine-Grained Spatiotemporal Understanding in LMMsHanyu Zhou, Gim Hee LeeICCV 2025 · 被引用 6 次
