Egocentric Auditory Attention Localization in Conversations
Fiona Ryan, Hao Jiang, Abhinav Shukla, James M. Rehg, Vamsi Krishna Ithapu
Abstract
In a noisy conversation environment such as a dinner party, people often exhibit selective auditory attention, or the ability to focus on a particular speaker while tuning out others. Recognizing who somebody is listening to in a conversation is essential for developing technologies that can understand social behavior and devices that can augment human hearing by amplifying particular sound sources. The computer vision and audio research communities have made great strides towards recognizing sound sources and speakers in scenes. In this work, we take a step further by focusing on the problem of localizing auditory attention targets in egocentric video, or detecting who in a camera wearer's field of view they are listening to. To tackle the new and challenging Selective Auditory Attention Localization problem, we propose an end-to-end deep learning approach that uses egocentric video and multichannel audio to predict the heatmap of the camera wearer's auditory attention. Our approach leverages spatiotemporal audiovisual features and holistic reasoning about the scene to make predictions, and outperforms a set of baselines on a challenging multi-speaker conversation dataset. Project
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2cf59d2b-223b-4123-96e9-6f136f8ecdd2Cited by top-tier papers9
- Multi-speaker Attention Alignment for Multimodal Social InteractionLiangyang Ouyang, Yifei Huang, Mingfang Zhang, Caixin Kang et al.CVPR 2026 · 8 citations
- EgoDTM: Towards 3D-Aware Egocentric Video-Language PretrainingBoshen Xu, Yuting Mei, Xinbi Liu, Sipeng Zheng et al.NeurIPS 2025 · 6 citations
- Learning Spatial Features from Audio-Visual Correspondence in Egocentric VideosSagnik Majumder, Ziad Al-Halah, Kristen GraumanCVPR 2024 · 3 citations
- Omni-MMSI: Toward Identity-attributed Social Interaction UnderstandingXinpeng Li, Bolin Lai, Hardy Chen, Shijian Deng et al.CVPR 2026 · 3 citations
- EAGLE: Egocentric AGgregated Language-video EngineJing Bi, Yunlong Tang, Luchuan Song, Ali Vosoughi et al.ACM MM 2024 · 3 citations
Builds on14
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Attention Bottlenecks for Multimodal FusionArsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen et al.NeurIPS 2021 · 884 citations
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and TextHassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang et al.NeurIPS 2021 · 782 citations
- Self-Supervised Learning by Cross-Modal Audio-Video ClusteringHumam Alwassel, Dhruv Mahajan, Bruno Korbar, Lorenzo Torresani et al.NeurIPS 2020 · 483 citations
Related papers
- The Audio-Visual Conversational Graph: From an Egocentric-Exocentric PerspectiveWenqi Jia, Miao Liu, Hao Jiang, Ishwarya Ananthabhotla et al.CVPR 2024
- Egocentric Deep Multi-Channel Audio-Visual Active Speaker LocalizationHao Jiang, Calvin Murdock, Vamsi Krishna IthapuCVPR 2022 · 39 citations
- Seeing Conversations: Communication Context Identification in Egocentric VideoTobias Dorszewski, Jens HjortkjærCVPR 2026
- Clink! Chop! Thud! - Learning Object Sounds From Real-World InteractionsMengyu Yang, Yiming Chen, Haozheng Pei, Siddhant Agarwal et al.ICCV 2025
- Detecting Attended Visual Targets in VideoEunji Chong, Yongxin Wang, Nataniel Ruiz, James M. RehgCVPR 2020
