UniCon: Unified Context Network for Robust Active Speaker Detection
Yuanhang Zhang, Susan Liang, Shuang Yang, Xiao Liu, Zhongqin Wu, Shiguang Shan, Xilin Chen
Abstract
We propose a new efficient framework, the Unified Context Network (UniCon), for robust active speaker detection (ASD). Traditional methods for ASD usually operate on each candidate's pre-cropped face track separately and do not sufficiently consider the relationships among the candidates. This potentially limits performance, especially in challenging scenarios with low-resolution faces, multiple candidates, etc. Our solution is a novel, unified framework that focuses on jointly modeling multiple types of contextual information: spatial context to indicate the position and scale of each candidate's face, relational context to capture the visual relationships among the candidates and contrast audio-visual affinities with each other, and temporal context to aggregate long-term information and smooth out local uncertainties. Based on such information, our model optimizes all candidates in a unified process for robust and reliable ASD. A thorough ablation study is performed on several challenging ASD benchmarks under different settings. In particular, our method outperforms the state-of-the-art by a large margin of about 15% mean Average Precision (mAP) absolute on two challenging subsets: one with three candidate speakers, and the other with faces smaller than 64 pixels. Together, our UniCon achieves 92.0% mAP on the AVA-ActiveSpeaker validation set, surpassing 90% for the first time on this challenging dataset at the time of submission. Project website: https://unicon-asd.github.io/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext be105ce9-b047-4fdd-b972-57f369c8f309Cited by top-tier papers6
- AV-NeRF: Learning Neural Fields for Real-World Audio-Visual Scene SynthesisSusan Liang, Chao Huang, Yapeng Tian, Anurag Kumar et al.NeurIPS 2023 · 77 citations
- LoCoNet: Long-Short Context Network for Active Speaker DetectionXizi Wang, Feng Cheng, Gedas BertasiusCVPR 2024 · 25 citations
- Learning Spatial Features from Audio-Visual Correspondence in Egocentric VideosSagnik Majumder, Ziad Al-Halah, Kristen GraumanCVPR 2024 · 3 citations
- A Light Weight Model for Active Speaker DetectionJunhua Liao, Haihan Duan, Kanghui Feng, Wanbing Zhao et al.CVPR 2023
- BinauralFlow: A Causal and Streamable Approach for High-Quality Binaural Speech Synthesis with Flow Matching ModelsSusan Liang, Dejan Markovic, Israel D. Gebru, Steven Krenn et al.ICML 2025
Builds on3
- Actor-Context-Actor Relation Network for Spatio-Temporal Action LocalizationJunting Pan, Siyu Chen, Mike Zheng Shou, Yu Liu et al.CVPR 2021
- RetinaFace: Single-Shot Multi-Level Face Localisation in the WildJiankang Deng, Jia Guo, Evangelos Ververas, Irene Kotsia et al.CVPR 2020
- Active Speakers in ContextJuan León Alcázar, Fabian Caba, Long Mai, Federico Perazzi et al.CVPR 2020
Related papers
- How to Design a Three-Stage Architecture for Audio-Visual Active Speaker Detection in the WildOkan Köpüklü, Maja Taseska, Gerhard RigollICCV 2021 · 59 citations
- Is Someone Speaking?: Exploring Long-term Temporal Features for Audio-visual Active Speaker DetectionRuijie Tao, Zexu Pan, Rohan Kumar Das, Xinyuan Qian et al.ACM MM 2021 · 154 citations
- MAAS: Multi-modal Assignation for Active Speaker DetectionJuan León Alcázar, Fabian Caba Heilbron, Ali K. Thabet, Bernard GhanemICCV 2021 · 66 citations
- AVA-AVD: Audio-visual Speaker Diarization in the WildEric Zhongcong Xu, Zeyang Song, Satoshi Tsutsui, Chao Feng et al.ACM MM 2022 · 34 citations
- AVTrack: Audio-Visual Tracking in Human-centric Complex ScenesYaoting Wang, Yun Zhou, Zipei Zhang, Henghui DingICML 2026 · 1 citation
