MAAS: Multi-modal Assignation for Active Speaker Detection
Juan León Alcázar, Fabian Caba Heilbron, Ali K. Thabet, Bernard Ghanem
摘要
Active speaker detection requires a mindful integration of multi-modal cues. Current methods focus on modeling and fusing short-term audiovisual features for individual speakers, often at frame level. We present a novel approach to active speaker detection that directly addresses the multimodal nature of the problem and provides a straightforward strategy, where independent visual features (speakers) in the scene are assigned to a previously detected speech event. Our experiments show that a small graph data structure built from local information can approximate an instantaneous audio-visual assignment problem. Moreover, the temporal extension of this initial graph achieves a new state-of-the-art performance on the AVA-ActiveSpeaker dataset with a mAP of 88.8%. We have made our code available at https://github.com/fuankarion/MAAS .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Is Someone Speaking?: Exploring Long-term Temporal Features for Audio-visual Active Speaker DetectionRuijie Tao, Zexu Pan, Rohan Kumar Das, Xinyuan Qian 等ACM MM 2021 · 被引用 154 次
- MIntRec: A New Dataset for Multimodal Intent RecognitionHanlei Zhang, Hua Xu, Xin Wang, Qianrui Zhou 等ACM MM 2022 · 被引用 66 次
- PEANUT: A Human-AI Collaborative Tool for Annotating Audio-Visual DataZheng Zhang, Zheng Ning, Chenliang Xu, Yapeng Tian 等UIST 2023 · 被引用 9 次
它引用的顶会 Paper3
- DeepGCNs: Can GCNs Go As Deep As CNNs?Guohao Li, Matthias Müller, Ali K. Thabet, Bernard GhanemICCV 2019 · 被引用 1,586 次
- G-TAD: Sub-Graph Localization for Temporal Action DetectionMengmeng Xu, Chen Zhao, David S. Rojas, Ali K. Thabet 等CVPR 2020
- Active Speakers in ContextJuan León Alcázar, Fabian Caba, Long Mai, Federico Perazzi 等CVPR 2020
相关 Paper
- A Light Weight Model for Active Speaker DetectionJunhua Liao, Haihan Duan, Kanghui Feng, Wanbing Zhao 等CVPR 2023
- How to Design a Three-Stage Architecture for Audio-Visual Active Speaker Detection in the WildOkan Köpüklü, Maja Taseska, Gerhard RigollICCV 2021 · 被引用 59 次
- LoCoNet: Long-Short Context Network for Active Speaker DetectionXizi Wang, Feng Cheng, Gedas BertasiusCVPR 2024 · 被引用 25 次
- UniCon: Unified Context Network for Robust Active Speaker DetectionYuanhang Zhang, Susan Liang, Shuang Yang, Xiao Liu 等ACM MM 2021 · 被引用 40 次
- Integrating Audio, Visual, and Semantic Information for Enhanced Multimodal Speaker Diarization on Multi-party ConversationLuyao Cheng, Hui Wang, Chong Deng, Siqi Zheng 等ACL 2025
