MTGA: Multi-View Temporal Granularity Aligned Aggregation for Event-Based Lip-Reading
Wenhao Zhang, Jun Wang, Yong Luo, Lei Yu, Wei Yu, Zheng He, Jialie Shen
Abstract
Lip-reading is to utilize the visual information of the speaker’s lip movements to recognize words and sentences. Existing event-based lip-reading solutions integrate different frame rate branches to learn spatio-temporal features of varying granularities. However, aggregating events into event frames inevitably leads to the loss of fine-grained temporal information within frames. To remedy this drawback, we propose a novel framework termed Multi-view Temporal Granularity aligned Aggregation (MTGA). Specifically, we first present a novel event representation method, namely time-segmented voxel graph list, where the most significant local voxels are temporally connected into a graph list. Then we design a spatio-temporal fusion module based on temporal granularity alignment, where the global spatial features extracted from event frames, together with the local relative spatial and temporal features contained in voxel graph list are effectively aligned and integrated. Finally, we design a temporal aggregation module that incorporates positional encoding, which enables the capture of local absolute spatial and global temporal information. Experiments demonstrate that our method outperforms both the event-based and video-based lip-reading counterparts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ba0f1804-b831-4762-a622-bc3f890ec7edBuilds on6
- End-to-End Learning of Representations for Asynchronous Event-Based DataDaniel Gehrig, Antonio Loquercio, Konstantinos G. Derpanis, Davide ScaramuzzaICCV 2019 · 427 citations
- Sub-word Level Lip Reading With Visual AttentionK. R. Prajwal, Triantafyllos Afouras, Andrew ZissermanCVPR 2022 · 104 citations
- Spatio-Temporal Fusion Based Convolutional Sequence Learning for Lip ReadingXingxuan Zhang, Feng Cheng, Shilin WangICCV 2019 · 87 citations
- GET: Group Event Transformer for Event-Based VisionYansong Peng, Yueyi Zhang, Zhiwei Xiong, Xiaoyan Sun et al.ICCV 2023 · 86 citations
- Multi-grained Spatio-Temporal Features Perceived Network for Event-based Lip-ReadingGanchao Tan, Yang Wang, Han Han, Yang Cao et al.CVPR 2022 · 36 citations
Related papers
- EventLip: Enhancing Event-Based Lip Reading via Frequency-Aware Spatiotemporal Hypergraph ModelingXueyi Zhang, Jialu Sun, Chengwei Zhang, Xianghu Yue et al.ACM MM 2025
- T2SGrid: Temporal-to-Spatial Gridification for Video Temporal GroundingChaohong Guo, Yihan He, Yongwei Nie, Fei Ma et al.CVPR 2026 · 2 citations
- LLaFEA: Frame-Event Complementary Fusion for Fine-Grained Spatiotemporal Understanding in LMMsHanyu Zhou, Gim Hee LeeICCV 2025 · 6 citations
- Multi-Level Representation Learning with Semantic Alignment for Referring Video Object SegmentationDongming Wu, Xingping Dong, Ling Shao, Jianbing ShenCVPR 2022 · 55 citations
- Multi-Granularity Reference-Aided Attentive Feature Aggregation for Video-Based Person Re-IdentificationZhizheng Zhang, Cuiling Lan, Wenjun Zeng, Zhibo ChenCVPR 2020
