EgoSound: Benchmarking Sound Understanding in Egocentric Videos
Bingwen Zhu, Yuqian Fu, Qiaole Dong, Guolei Sun, Tianwen Qian, Yuzheng Wu, Danda Pani Paudel, Yanwei Fu, Xiangyang Xue
摘要
Multimodal Large Language Models (MLLMs) have recently achieved remarkable progress in vision-language understanding. Yet, human perception is inherently multisensory, integrating sight, sound, and motion to reason about the world. Among these modalities, sound provides indispensable cues about spatial layout, off-screen events, and causal interactions, particularly in egocentric settings where auditory and visual signals are tightly coupled. To this end, we introduce EgoSound, the first benchmark designed to systematically evaluate egocentric sound understanding in MLLMs. EgoSound unifies data from Ego4D and EgoBlind, encompassing both sighted and sound-dependent experiences. It defines a seven-task taxonomy spanning intrinsic sound perception, spatial localization, causal inference, and cross-modal reasoning. Constructed through a multi-stage auto-generative pipeline, EgoSound contains 7315 validated QA pairs across 900 videos. Comprehensive experiments on nine state-of-the-art MLLMs reveal that current models exhibit emerging auditory reasoning abilities but remain limited in fine-grained spatial and causal understanding. EgoSound establishes a challenging foundation for advancing multisensory egocentric intelligence, bridging the gap between seeing and truly hearing the world. Project page: https://groolegend.github.io/EgoSound/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper23
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis 等CVPR 2022 · 被引用 525 次
- EgoVLPv2: Egocentric Video-Language Pre-training with Fusion in the BackboneShraman Pramanick, Yale Song, Sayan Nag, Kevin Qinghong Lin 等ICCV 2023 · 被引用 152 次
- HoloAssist: an Egocentric Human Interaction Dataset for Interactive AI Assistants in the Real WorldXin Wang, Taein Kwon, Mahdi Rad, Bowen Pan 等ICCV 2023 · 被引用 151 次
- BAT: Learning to Reason about Spatial Sounds with Large Language ModelsZhisheng Zheng, Puyuan Peng, Ziyang Ma, Xie Chen 等ICML 2024 · 被引用 42 次
- Omni-Captioner: Data Pipeline, Models, and Benchmark for Omni Detailed PerceptionZiyang Ma, Ruiyang Xu, Zhenghao Xing, Yunfei Chu 等ICLR 2026 · 被引用 36 次
相关 Paper
- EgoAVU: Egocentric Audio-Visual UnderstandingAshish Seth, Xinhao Mei, Changsheng Zhao, Varun Nagaraja 等CVPR 2026 · 被引用 1 次
- OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMsCaorui Li, Yu Chen, Yiyan Ji, Jin Xu 等ICLR 2026 · 被引用 53 次
- AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMsYaoting Wang, Ziyi Zhang, Wenming Tu, Shaoxuan Xu 等ICML 2026
- EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive HierarchyJinzhao Li, Yinuo Chen, Dongxu Piao, Panwang Pan 等CVPR 2026 · 被引用 2 次
- EGOILLUSION: Benchmarking Hallucinations in Egocentric Video UnderstandingAshish Seth, Utkarsh Tyagi, Ramaneswaran Selvakumar, Nishit Anand 等EMNLP 2025 · 被引用 1 次
