MTGA: Multi-View Temporal Granularity Aligned Aggregation for Event-Based Lip-Reading
Wenhao Zhang, Jun Wang, Yong Luo, Lei Yu, Wei Yu, Zheng He, Jialie Shen
摘要
Lip-reading is to utilize the visual information of the speaker’s lip movements to recognize words and sentences. Existing event-based lip-reading solutions integrate different frame rate branches to learn spatio-temporal features of varying granularities. However, aggregating events into event frames inevitably leads to the loss of fine-grained temporal information within frames. To remedy this drawback, we propose a novel framework termed Multi-view Temporal Granularity aligned Aggregation (MTGA). Specifically, we first present a novel event representation method, namely time-segmented voxel graph list, where the most significant local voxels are temporally connected into a graph list. Then we design a spatio-temporal fusion module based on temporal granularity alignment, where the global spatial features extracted from event frames, together with the local relative spatial and temporal features contained in voxel graph list are effectively aligned and integrated. Finally, we design a temporal aggregation module that incorporates positional encoding, which enables the capture of local absolute spatial and global temporal information. Experiments demonstrate that our method outperforms both the event-based and video-based lip-reading counterparts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- End-to-End Learning of Representations for Asynchronous Event-Based DataDaniel Gehrig, Antonio Loquercio, Konstantinos G. Derpanis, Davide ScaramuzzaICCV 2019 · 被引用 427 次
- Sub-word Level Lip Reading With Visual AttentionK. R. Prajwal, Triantafyllos Afouras, Andrew ZissermanCVPR 2022 · 被引用 104 次
- Spatio-Temporal Fusion Based Convolutional Sequence Learning for Lip ReadingXingxuan Zhang, Feng Cheng, Shilin WangICCV 2019 · 被引用 87 次
- GET: Group Event Transformer for Event-Based VisionYansong Peng, Yueyi Zhang, Zhiwei Xiong, Xiaoyan Sun 等ICCV 2023 · 被引用 86 次
- Multi-grained Spatio-Temporal Features Perceived Network for Event-based Lip-ReadingGanchao Tan, Yang Wang, Han Han, Yang Cao 等CVPR 2022 · 被引用 36 次
相关 Paper
- EventLip: Enhancing Event-Based Lip Reading via Frequency-Aware Spatiotemporal Hypergraph ModelingXueyi Zhang, Jialu Sun, Chengwei Zhang, Xianghu Yue 等ACM MM 2025
- T2SGrid: Temporal-to-Spatial Gridification for Video Temporal GroundingChaohong Guo, Yihan He, Yongwei Nie, Fei Ma 等CVPR 2026 · 被引用 2 次
- LLaFEA: Frame-Event Complementary Fusion for Fine-Grained Spatiotemporal Understanding in LMMsHanyu Zhou, Gim Hee LeeICCV 2025 · 被引用 6 次
- Multi-Level Representation Learning with Semantic Alignment for Referring Video Object SegmentationDongming Wu, Xingping Dong, Ling Shao, Jianbing ShenCVPR 2022 · 被引用 55 次
- Multi-Granularity Reference-Aided Attentive Feature Aggregation for Video-Based Person Re-IdentificationZhizheng Zhang, Cuiling Lan, Wenjun Zeng, Zhibo ChenCVPR 2020
