LoCoNet: Long-Short Context Network for Active Speaker Detection
Xizi Wang, Feng Cheng, Gedas Bertasius
Abstract
Active Speaker Detection (ASD) aims to identify who is speaking in each frame of a video. Solving ASD involves using audio and visual information in two complementary contexts: long-term intra-speaker context models the temporal dependencies of the same speaker, while short-term inter-speaker context models the interactions of speakers in the same scene. Motivated by these observations, we propose LoCoNet, a simple but effective Long-Short Context Network that leverages Long-term Intra-speaker Modeling (LIM) and Short-term Inter-speaker Modeling (SIM) in an interleaved manner. LIM employs self-attention for long-range temporal dependencies modeling and cross-attention for audio-visual interactions modeling. SIM incorporates convolutional blocks that capture local patterns for short-term inter-speaker context. Experiments show that LoCoNet achieves state-of-the-art performance on multiple datasets, with 95.2% (+0.3%) mAP on AVA-ActiveSpeaker, 97.2% (+2.7%) mAP on Talkies, and 68.4% (+7.7%) mAP on Ego4D. Moreover, in challenging cases where multiple speakers are present, LoCoNet outperforms previous state-of-the-art methods by 3.0% mAP on AVA-ActiveSpeaker. The code is available at https://github.com/SJTUwxz/LoCoNet_ASD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4c57560d-9e8c-4e43-b7f8-7efbbb23b0a5Cited by top-tier papers2
- A Light Weight Model for Active Speaker DetectionJunhua Liao, Haihan Duan, Kanghui Feng, Wanbing Zhao et al.CVPR 2023
- EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric PerceptionSanjoy Chowdhury, Subrata Biswas, Sayan Nag, Tushar Nagarajan et al.ICCV 2025
Builds on11
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
- EPIC-Fusion: Audio-Visual Temporal Binding for Egocentric Action RecognitionEvangelos Kazakos, Arsha Nagrani, Andrew Zisserman, Dima DamenICCV 2019 · 395 citations
- Is Someone Speaking?: Exploring Long-term Temporal Features for Audio-visual Active Speaker DetectionRuijie Tao, Zexu Pan, Rohan Kumar Das, Xinyuan Qian et al.ACM MM 2021 · 154 citations
- How to Design a Three-Stage Architecture for Audio-Visual Active Speaker Detection in the WildOkan Köpüklü, Maja Taseska, Gerhard RigollICCV 2021 · 59 citations
- UniCon: Unified Context Network for Robust Active Speaker DetectionYuanhang Zhang, Susan Liang, Shuang Yang, Xiao Liu et al.ACM MM 2021 · 40 citations
Related papers
- Active Speakers in ContextJuan León Alcázar, Fabian Caba, Long Mai, Federico Perazzi et al.CVPR 2020
- MAAS: Multi-modal Assignation for Active Speaker DetectionJuan León Alcázar, Fabian Caba Heilbron, Ali K. Thabet, Bernard GhanemICCV 2021 · 66 citations
- AVA-AVD: Audio-visual Speaker Diarization in the WildEric Zhongcong Xu, Zeyang Song, Satoshi Tsutsui, Chao Feng et al.ACM MM 2022 · 34 citations
- ACSNet: Action-Context Separation Network for Weakly Supervised Temporal Action LocalizationZiyi Liu, Le Wang, Qilin Zhang, Wei Tang et al.AAAI 2021 · 83 citations
- AVTrack: Audio-Visual Tracking in Human-centric Complex ScenesYaoting Wang, Yun Zhou, Zipei Zhang, Henghui DingICML 2026 · 1 citation
