Multi-grained Spatio-Temporal Features Perceived Network for Event-based Lip-Reading
Ganchao Tan, Yang Wang, Han Han, Yang Cao, Feng Wu, Zhengjun Zha
Abstract
Automatic lip-reading (ALR) aims to recognize words using visual information from the speaker's lip movements. In this work, we introduce a novel type of sensing device, event cameras, for the task of ALR. Event cameras have both technical and application advantages over conventional cameras for the ALR task because they have higher temporal resolution, less redundant visual information, and lower power consumption. To recognize words from the event data, we propose a novel Multi-grained Spatio-Temporal Features Perceived Network (MSTP) to perceive fine-grained spatio-temporal features from microsecond time-resolved event data. Specifically, a multi-branch network architecture is designed, in which different grained spatio-temporal features are learned by operating at different frame rates. The branch operating on the low frame rate can perceive spatial complete but temporal coarse features. While the branch operating on the high frame rate can perceive spatial coarse but temporal refinement features. And a message flow module is devised to integrate the features from different branches, leading to perceiving more discriminative spatio-temporal features. In addition, we present the first event-based lip-reading dataset (DVS-Lip) captured by the event camera. Experimental results demonstrated the superiority of the proposed model compared to the state-of-the-art event-based action recognition models and video-based lip-reading models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4b4243c5-74cc-4ea3-b6f8-52a53cb3a54cCited by top-tier papers8
- Dilated convolution with learnable spacingsIsmail Khalfaoui Hassani, Thomas Pellegrini, Timothée MasquelierICLR 2023 · 18 citations
- MTGA: Multi-View Temporal Granularity Aligned Aggregation for Event-Based Lip-ReadingWenhao Zhang, Jun Wang, Yong Luo, Lei Yu et al.AAAI 2025 · 8 citations
- Fractional-Order Spiking Neural NetworkChengjie Ge, Yufeng Peng, Zihao Li, Qiyu Kang et al.ICLR 2026 · 5 citations
- Multiplication-Free Parallelizable Spiking Neurons with Efficient Spatio-Temporal DynamicsPeng Xue, Wei Fang, Zhengyu Ma, Zihan Huang et al.NeurIPS 2025 · 5 citations
- Event2Vec: Processing neuromorphic events directly by representations in vector spaceWei Fang, Priyadarshini PandaICML 2026 · 4 citations
Builds on8
- ViTAE: Vision Transformer Advanced by Exploring Intrinsic Inductive BiasYufei Xu, Qiming Zhang, Jing Zhang, Dacheng TaoNeurIPS 2021 · 429 citations
- End-to-End Learning of Representations for Asynchronous Event-Based DataDaniel Gehrig, Antonio Loquercio, Konstantinos G. Derpanis, Davide ScaramuzzaICCV 2019 · 427 citations
- DualLip: A System for Joint Lip Reading and GenerationWeicong Chen, Xu Tan, Yingce Xia, Tao Qin et al.ACM MM 2020 · 25 citations
- SimulLR: Simultaneous Lip Reading Transducer with Attention-Guided Adaptive MemoryZhijie Lin, Zhou Zhao, Haoyuan Li, Jinglin Liu et al.ACM MM 2021 · 13 citations
- Time Lens: Event-Based Video Frame InterpolationStepan Tulyakov, Daniel Gehrig, Stamatios Georgoulis, Julius Erbach et al.CVPR 2021
Related papers
- EventLip: Enhancing Event-Based Lip Reading via Frequency-Aware Spatiotemporal Hypergraph ModelingXueyi Zhang, Jialu Sun, Chengwei Zhang, Xianghu Yue et al.ACM MM 2025
- Discriminative Multi-Modality Speech RecognitionBo Xu, Cheng Lu, Yandong Guo, Jacob WangCVPR 2020
- Hearing Lips: Improving Lip Reading by Distilling Speech RecognizersYa Zhao, Rui Xu, Xinchao Wang, Peng Hou et al.AAAI 2020 · 106 citations
- Learning an Event Sequence Embedding for Dense Event-Based Deep StereoStepan Tulyakov, François Fleuret, Martin Kiefel, Peter V. Gehler et al.ICCV 2019 · 122 citations
- LLaFEA: Frame-Event Complementary Fusion for Fine-Grained Spatiotemporal Understanding in LMMsHanyu Zhou, Gim Hee LeeICCV 2025 · 6 citations
