BAT: Learning to Reason about Spatial Sounds with Large Language Models
Zhisheng Zheng, Puyuan Peng, Ziyang Ma, Xie Chen, Eunsol Choi, David Harwath
摘要
Spatial sound reasoning is a fundamental human skill, enabling us to navigate and interpret our surroundings based on sound. In this paper we present BAT, which combines the spatial sound perception ability of a binaural acoustic scene analysis model with the natural language reasoning capabilities of a large language model (LLM) to replicate this innate ability. To address the lack of existing datasets of in-the-wild spatial sounds, we synthesized a binaural audio dataset using Au-dioSet and SoundSpaces 2.0. Next, we developed SPATIALSOUNDQA, a spatial sound-based question-answering dataset, offering a range of QA tasks that train BAT in various aspects of spatial sound perception and reasoning. The acoustic front end encoder of BAT is a novel spatial audio encoder named Spatial Audio Spectrogram Transformer, or SPATIAL-AST, which by itself achieves strong performance across sound event detection, spatial localization, and distance estimation. By integrating SPATIAL-AST with LLaMA-2 7B model, BAT transcends standard Sound Event Localization and Detection (SELD) tasks, enabling the model to reason about the relationships between the sounds in its environment. Our experiments demonstrate BAT's superior performance on both spatial sound perception and reasoning, showcasing the immense potential of LLMs in navigating and interpreting complex spatial audio environments. Our demo, dataset, code and model weights are available at: https://zhishengzheng.com/bat .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Learning Spatially-Aware Language and Audio EmbeddingsBhavika Devnani, Skyler Seto, Zakaria Aldeneh, Alessandro Toso 等NeurIPS 2024 · 被引用 31 次
- SPHERE: Unveiling Spatial Blind Spots in Vision-Language Models Through Hierarchical EvaluationWenyu Zhang, Wei En Ng, Lixin Ma, Yuwen Wang 等ACL 2025 · 被引用 20 次
- SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and HearingMingfei Chen, Zijun Cui, Xiulong Liu, Jinlin Xiang 等NeurIPS 2025 · 被引用 18 次
- 🎧MOSPA: Human Motion Generation Driven by Spatial AudioShuyang Xu, Zhiyang Dou, Mingyi Shi, Liang Pan 等NeurIPS 2025 · 被引用 13 次
- OWL : Geometry-Aware Spatial Reasoning for Audio Large Language ModelsSubrata Biswas, Mohammad Nur Hossain Khan, Bashima IslamICLR 2026 · 被引用 9 次
它引用的顶会 Paper10
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Grounding Multimodal Large Language Models to the WorldZhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao 等ICLR 2024 · 被引用 1,170 次
相关 Paper
- Hear you are: Teaching LLMs Spatial Reasoning with Vision and Spatial SoundHyeonggon Ryu, Joon Son Chung, David HarwathCVPR 2026 · 被引用 4 次
- Listen, Think, and UnderstandYuan Gong, Hongyin Luo, Alexander H. Liu, Leonid Karlinsky 等ICLR 2024 · 被引用 247 次
- R-AVST: Empowering Video-LLMs with Fine-Grained Spatio-Temporal Reasoning in Complex Audio-Visual ScenariosLu Zhu, Tiantian Geng, Yangye Chen, Teng Wang 等AAAI 2026 · 被引用 1 次
- An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal ModelsFatemeh Shiri, Xiao-Yu Guo, Mona Far, Xin Yu 等EMNLP 2024 · 被引用 7 次
- In-the-wild Audio Spatialization with Flexible Text-guided LocalizationTianrui Pan, Jie Liu, Zewen Huang, Jie Tang 等ACL 2025 · 被引用 2 次
