How to Design a Three-Stage Architecture for Audio-Visual Active Speaker Detection in the Wild
Okan Köpüklü, Maja Taseska, Gerhard Rigoll
Abstract
Successful active speaker detection requires a threestage pipeline: (i) audio-visual encoding for all speakers in the clip, (ii) inter-speaker relation modeling between a reference speaker and the background speakers within each frame, and (iii) temporal modeling for the reference speaker. Each stage of this pipeline plays an important role for the final performance of the created architecture. Based on a series of controlled experiments, this work presents several practical guidelines for audio-visual active speaker detection. Correspondingly, we present a new architecture called ASDNet, which achieves a new state-of-the-art on the AVA-ActiveSpeaker dataset with a mAP of 93.5% outperforming the second best with a large margin of 4.7%. Our code and pretrained models are publicly available 1 . Recently, the AVA-ActiveSpeaker dataset [42] provided the first large-scale standard benchmark for audio-visual active speaker detection in the wild. Recent research [1, 32] indicates that active speaker detection in the wild requires (i) integration of audio-visual information for each speaker, (ii) contextual information that captures inter-speaker relationships, and (iii) temporal modeling to exploit long term relationships in natural conversation. In this paper, we con-
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a4dbaab5-29bb-4167-98e0-d670995474d0Cited by top-tier papers6
- Egocentric Deep Multi-Channel Audio-Visual Active Speaker LocalizationHao Jiang, Calvin Murdock, Vamsi Krishna IthapuCVPR 2022 · 39 citations
- LoCoNet: Long-Short Context Network for Active Speaker DetectionXizi Wang, Feng Cheng, Gedas BertasiusCVPR 2024 · 25 citations
- Learning Spatial Features from Audio-Visual Correspondence in Egocentric VideosSagnik Majumder, Ziad Al-Halah, Kristen GraumanCVPR 2024 · 3 citations
- A Light Weight Model for Active Speaker DetectionJunhua Liao, Haihan Duan, Kanghui Feng, Wanbing Zhao et al.CVPR 2023
- Egocentric Auditory Attention Localization in ConversationsFiona Ryan, Hao Jiang, Abhinav Shukla, James M. Rehg et al.CVPR 2023
Builds on4
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Listen to Look: Action Recognition by Previewing AudioRuohan Gao, Tae-Hyun Oh, Kristen Grauman, Lorenzo TorresaniCVPR 2020
- Active Speakers in ContextJuan León Alcázar, Fabian Caba, Long Mai, Federico Perazzi et al.CVPR 2020
- Speech2Action: Cross-Modal Supervision for Action RecognitionArsha Nagrani, Chen Sun, David Ross, Rahul Sukthankar et al.CVPR 2020
Related papers
- Is Someone Speaking?: Exploring Long-term Temporal Features for Audio-visual Active Speaker DetectionRuijie Tao, Zexu Pan, Rohan Kumar Das, Xinyuan Qian et al.ACM MM 2021 · 154 citations
- MAAS: Multi-modal Assignation for Active Speaker DetectionJuan León Alcázar, Fabian Caba Heilbron, Ali K. Thabet, Bernard GhanemICCV 2021 · 66 citations
- AVA-AVD: Audio-visual Speaker Diarization in the WildEric Zhongcong Xu, Zeyang Song, Satoshi Tsutsui, Chao Feng et al.ACM MM 2022 · 34 citations
- UniCon: Unified Context Network for Robust Active Speaker DetectionYuanhang Zhang, Susan Liang, Shuang Yang, Xiao Liu et al.ACM MM 2021 · 40 citations
- The Right to Talk: An Audio-Visual Transformer ApproachThanh-Dat Truong, Chi Nhan Duong, The De Vu, Hoang Anh Pham et al.ICCV 2021 · 39 citations
