Physics-Informed Audio-Geometry-Grid Representation Learning for Universal Sound Source Localization
Min-Sang Baek, Gyeong-Su Kim, Donghyun Kim, Joon-Hyuk Chang
摘要
Sound source localization (SSL) is a fundamental task in spatial audio understanding, yet most deep neural network-based methods are constrained by fixed array geometries and predefined directional grids, limiting generalizability and scalability. To address these issues, we propose audio-geometry-grid representation learning (AGG-RL), a novel framework that jointly learns audio-geometry and grid representations in a shared latent space, enabling both geometry-invariant and grid-flexible SSL. Moreover, to enhance generalizability and interpretability, we introduce two physics-informed components: a learnable non-uniform discrete Fourier transform (LNuDFT), which optimizes the dense allocation of frequency bins in a non-uniform manner to emphasize informative phase regions, and a relative microphone positional encoding (rMPE), which encodes relative microphone coordinates in accordance with the nature of inter-channel time differences. Experiments on synthetic and real datasets demonstrate that AGG-RL achieves superior performance, particularly under unseen conditions. The results highlight the potential of representation learning with physics-informed design towards a universal solution for spatial acoustic scene understanding across diverse scenarios.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- Learning Audio-Visual Speech Representation by Masked Multimodal Cluster PredictionBowen Shi, Wei-Ning Hsu, Kushal Lakhotia, Abdelrahman MohamedICLR 2022 · 被引用 460 次
- The Impact of Positional Encoding on Length Generalization in TransformersAmirhossein Kazemnejad, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Payel Das 等NeurIPS 2023 · 被引用 444 次
- Learning Neural Acoustic FieldsAndrew F. Luo, Yilun Du, Michael J. Tarr, Josh Tenenbaum 等NeurIPS 2022 · 被引用 153 次
- BinauralGrad: A Two-Stage Conditional Diffusion Probabilistic Model for Binaural Audio SynthesisYichong Leng, Zehua Chen, Junliang Guo, Haohe Liu 等NeurIPS 2022 · 被引用 86 次
- Neural Synthesis of Binaural Speech From Mono AudioAlexander Richard, Dejan Markovic, Israel D. Gebru, Steven Krenn 等ICLR 2021 · 被引用 73 次
相关 Paper
- PhaseCoder: Microphone Geometry-Agnostic Spatial Audio Understanding for Multimodal LLMsArtem Dementyev, Wazeer Zulfikar, Sinan Hersek, Pascal Getreuer 等ICML 2026 · 被引用 4 次
- Exploiting Transformation Invariance and Equivariance for Self-supervised Sound LocalisationJinxiang Liu, Chen Ju, Weidi Xie, Ya ZhangACM MM 2022 · 被引用 37 次
- Self-Supervised Learning of Representations for Space Generates Multi-Modular Grid CellsRylan Schaeffer, Mikail Khona, Tzuhsuan Ma, Cristóbal Eyzaguirre 等NeurIPS 2023 · 被引用 40 次
- INRAS: Implicit Neural Representation for Audio ScenesKun Su, Mingfei Chen, Eli ShlizermanNeurIPS 2022 · 被引用 92 次
- Radiance-Field Guided Pretraining: Scaling Localization Models with Unlabeled Wireless SignalsGuosheng Wang, Shen Wang, Lei YangUbiComp 2026
