Is Someone Speaking?: Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection
Ruijie Tao, Zexu Pan, Rohan Kumar Das, Xinyuan Qian, Mike Zheng Shou, Haizhou Li
摘要
Active speaker detection (ASD) seeks to detect who is speaking in a visual scene of one or more speakers. The successful ASD depends on accurate interpretation of short-term and long-term audio and visual information, as well as audio-visual interaction. Unlike the prior work where systems make decision instantaneously using short-term features, we propose a novel framework, named TalkNet, that makes decision by taking both short-term and long-term features into consideration. TalkNet consists of audio and visual temporal encoders for feature representation, audio-visual cross-attention mechanism for inter-modality interaction, and a self-attention mechanism to capture long-term speaking evidence. The experiments demonstrate that TalkNet achieves 3.5% and 2.2% improvement over the state-of-the-art systems on the AVA-ActiveSpeaker dataset and Columbia ASD dataset, respectively. Code has been made available at: https://github.com/TaoRuijie/TalkNet_ASD.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper38
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis 等CVPR 2022 · 被引用 525 次
- Sub-word Level Lip Reading With Visual AttentionK. R. Prajwal, Triantafyllos Afouras, Andrew ZissermanCVPR 2022 · 被引用 104 次
- M3ED: Multi-modal Multi-scene Multi-label Emotional Dialogue DatabaseJinming Zhao, Tenggan Zhang, Jingwen Hu, Yuchen Liu 等ACL 2022 · 被引用 88 次
- MIntRec: A New Dataset for Multimodal Intent RecognitionHanlei Zhang, Hua Xu, Xin Wang, Qianrui Zhou 等ACM MM 2022 · 被引用 66 次
- A Facial Expression-Aware Multimodal Multi-task Learning Framework for Emotion Recognition in Multi-party ConversationsWenjie Zheng, Jianfei Yu, Rui Xia, Shijin WangACL 2023 · 被引用 38 次
它引用的顶会 Paper5
- Sequence Level Semantics Aggregation for Video Object DetectionHaiping Wu, Yuntao Chen, Naiyan Wang, Zhaoxiang ZhangICCV 2019 · 被引用 236 次
- Look, Listen, and Attend: Co-Attention Network for Self-Supervised Audio-Visual Representation LearningYing Cheng, Ruize Wang, Zhihao Pan, Rui Feng 等ACM MM 2020 · 被引用 93 次
- Leveraging Long-Range Temporal Relationships Between Proposals for Video Object DetectionMykhailo Shvets, Wei Liu, Alexander C. BergICCV 2019 · 被引用 91 次
- MAAS: Multi-modal Assignation for Active Speaker DetectionJuan León Alcázar, Fabian Caba Heilbron, Ali K. Thabet, Bernard GhanemICCV 2021 · 被引用 66 次
- Active Speakers in ContextJuan León Alcázar, Fabian Caba, Long Mai, Federico Perazzi 等CVPR 2020
相关 Paper
- LoCoNet: Long-Short Context Network for Active Speaker DetectionXizi Wang, Feng Cheng, Gedas BertasiusCVPR 2024 · 被引用 25 次
- How to Design a Three-Stage Architecture for Audio-Visual Active Speaker Detection in the WildOkan Köpüklü, Maja Taseska, Gerhard RigollICCV 2021 · 被引用 59 次
- A Light Weight Model for Active Speaker DetectionJunhua Liao, Haihan Duan, Kanghui Feng, Wanbing Zhao 等CVPR 2023
- UniCon: Unified Context Network for Robust Active Speaker DetectionYuanhang Zhang, Susan Liang, Shuang Yang, Xiao Liu 等ACM MM 2021 · 被引用 40 次
- AVA-AVD: Audio-visual Speaker Diarization in the WildEric Zhongcong Xu, Zeyang Song, Satoshi Tsutsui, Chao Feng 等ACM MM 2022 · 被引用 34 次
