How to Design a Three-Stage Architecture for Audio-Visual Active Speaker Detection in the Wild
Okan Köpüklü, Maja Taseska, Gerhard Rigoll
摘要
Successful active speaker detection requires a threestage pipeline: (i) audio-visual encoding for all speakers in the clip, (ii) inter-speaker relation modeling between a reference speaker and the background speakers within each frame, and (iii) temporal modeling for the reference speaker. Each stage of this pipeline plays an important role for the final performance of the created architecture. Based on a series of controlled experiments, this work presents several practical guidelines for audio-visual active speaker detection. Correspondingly, we present a new architecture called ASDNet, which achieves a new state-of-the-art on the AVA-ActiveSpeaker dataset with a mAP of 93.5% outperforming the second best with a large margin of 4.7%. Our code and pretrained models are publicly available 1 . Recently, the AVA-ActiveSpeaker dataset [42] provided the first large-scale standard benchmark for audio-visual active speaker detection in the wild. Recent research [1, 32] indicates that active speaker detection in the wild requires (i) integration of audio-visual information for each speaker, (ii) contextual information that captures inter-speaker relationships, and (iii) temporal modeling to exploit long term relationships in natural conversation. In this paper, we con-
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Egocentric Deep Multi-Channel Audio-Visual Active Speaker LocalizationHao Jiang, Calvin Murdock, Vamsi Krishna IthapuCVPR 2022 · 被引用 39 次
- LoCoNet: Long-Short Context Network for Active Speaker DetectionXizi Wang, Feng Cheng, Gedas BertasiusCVPR 2024 · 被引用 25 次
- Learning Spatial Features from Audio-Visual Correspondence in Egocentric VideosSagnik Majumder, Ziad Al-Halah, Kristen GraumanCVPR 2024 · 被引用 3 次
- A Light Weight Model for Active Speaker DetectionJunhua Liao, Haihan Duan, Kanghui Feng, Wanbing Zhao 等CVPR 2023
- Egocentric Auditory Attention Localization in ConversationsFiona Ryan, Hao Jiang, Abhinav Shukla, James M. Rehg 等CVPR 2023
它引用的顶会 Paper4
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- Listen to Look: Action Recognition by Previewing AudioRuohan Gao, Tae-Hyun Oh, Kristen Grauman, Lorenzo TorresaniCVPR 2020
- Active Speakers in ContextJuan León Alcázar, Fabian Caba, Long Mai, Federico Perazzi 等CVPR 2020
- Speech2Action: Cross-Modal Supervision for Action RecognitionArsha Nagrani, Chen Sun, David Ross, Rahul Sukthankar 等CVPR 2020
相关 Paper
- Is Someone Speaking?: Exploring Long-term Temporal Features for Audio-visual Active Speaker DetectionRuijie Tao, Zexu Pan, Rohan Kumar Das, Xinyuan Qian 等ACM MM 2021 · 被引用 154 次
- MAAS: Multi-modal Assignation for Active Speaker DetectionJuan León Alcázar, Fabian Caba Heilbron, Ali K. Thabet, Bernard GhanemICCV 2021 · 被引用 66 次
- AVA-AVD: Audio-visual Speaker Diarization in the WildEric Zhongcong Xu, Zeyang Song, Satoshi Tsutsui, Chao Feng 等ACM MM 2022 · 被引用 34 次
- UniCon: Unified Context Network for Robust Active Speaker DetectionYuanhang Zhang, Susan Liang, Shuang Yang, Xiao Liu 等ACM MM 2021 · 被引用 40 次
- The Right to Talk: An Audio-Visual Transformer ApproachThanh-Dat Truong, Chi Nhan Duong, The De Vu, Hoang Anh Pham 等ICCV 2021 · 被引用 39 次
