Egocentric Deep Multi-Channel Audio-Visual Active Speaker Localization
Hao Jiang, Calvin Murdock, Vamsi Krishna Ithapu
Abstract
Augmented reality devices have the potential to enhance human perception and enable other assistive functionalities in complex conversational environments. Effectively capturing the audio-visual context necessary for understanding these social interactions first requires detecting and localizing the voice activities of the device wearer and the surrounding people. These tasks are challenging due to their egocentric nature: the wearer's head motion may cause motion blur, surrounding people may appear in difficult viewing angles, and there may be occlusions, visual clutter, audio noise, and bad lighting. Under these conditions, previous state-of-the-art active speaker detection methods do not give satisfactory results. Instead, we tackle the problem from a new setting using both video and multi-channel microphone array audio. We propose a novel end-to-end deep learning approach that is able to give robust voice activity detection and localization results. In contrast to previous methods, our method localizes active speakers from all possible directions on the sphere, even outside the camera's field of view, while simultaneously detecting the device wearer's own voice activity. Our experiments show that the proposed method gives superior results, can run in real time, and is robust against noise and clutter.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bb46f828-bc11-4757-9e93-4340acc33e6bCited by top-tier papers12
- Few-Shot Audio-Visual Learning of Environment AcousticsSagnik Majumder, Changan Chen, Ziad Al-Halah, Kristen GraumanNeurIPS 2022 · 80 citations
- Sounding Bodies: Modeling 3D Spatial Sound of Humans Using Body Pose and AudioXudong Xu, Dejan Markovic, Jacob Sandakly, Todd Keebler et al.NeurIPS 2023 · 9 citations
- Learning Spatial Features from Audio-Visual Correspondence in Egocentric VideosSagnik Majumder, Ziad Al-Halah, Kristen GraumanCVPR 2024 · 3 citations
- EgoPPG: Heart Rate Estimation From Eye-Tracking Cameras in Egocentric Systems to Benefit Downstream Vision TasksBjörn Braun, Rayan Armani, Manuel Meier, Max Möbus et al.ICCV 2025 · 3 citations
- Proactive Hearing Assistants that Isolate Egocentric ConversationsGuilin Hu, Malek Itani, Tuochao Chen, Shyamnath GollakotaEMNLP 2025 · 1 citation
Builds on7
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- EPIC-Fusion: Audio-Visual Temporal Binding for Egocentric Action RecognitionEvangelos Kazakos, Arsha Nagrani, Andrew Zisserman, Dima DamenICCV 2019 · 395 citations
- Self-Supervised Moving Vehicle Tracking With Stereo SoundChuang Gan, Hang Zhao, Peihao Chen, David D. Cox et al.ICCV 2019 · 157 citations
- How to Design a Three-Stage Architecture for Audio-Visual Active Speaker Detection in the WildOkan Köpüklü, Maja Taseska, Gerhard RigollICCV 2021 · 59 citations
- The Right to Talk: An Audio-Visual Transformer ApproachThanh-Dat Truong, Chi Nhan Duong, The De Vu, Hoang Anh Pham et al.ICCV 2021 · 39 citations
Related papers
- Egocentric Auditory Attention Localization in ConversationsFiona Ryan, Hao Jiang, Abhinav Shukla, James M. Rehg et al.CVPR 2023
- Egocentric Pose Estimation from Human Vision SpanHao Jiang, Vamsi Krishna IthapuICCV 2021 · 37 citations
- Active Speakers in ContextJuan León Alcázar, Fabian Caba, Long Mai, Federico Perazzi et al.CVPR 2020
- UniCon: Unified Context Network for Robust Active Speaker DetectionYuanhang Zhang, Susan Liang, Shuang Yang, Xiao Liu et al.ACM MM 2021 · 40 citations
- MAAS: Multi-modal Assignation for Active Speaker DetectionJuan León Alcázar, Fabian Caba Heilbron, Ali K. Thabet, Bernard GhanemICCV 2021 · 66 citations
