Physics-Informed Audio-Geometry-Grid Representation Learning for Universal Sound Source Localization
Min-Sang Baek, Gyeong-Su Kim, Donghyun Kim, Joon-Hyuk Chang
Abstract
Sound source localization (SSL) is a fundamental task in spatial audio understanding, yet most deep neural network-based methods are constrained by fixed array geometries and predefined directional grids, limiting generalizability and scalability. To address these issues, we propose audio-geometry-grid representation learning (AGG-RL), a novel framework that jointly learns audio-geometry and grid representations in a shared latent space, enabling both geometry-invariant and grid-flexible SSL. Moreover, to enhance generalizability and interpretability, we introduce two physics-informed components: a learnable non-uniform discrete Fourier transform (LNuDFT), which optimizes the dense allocation of frequency bins in a non-uniform manner to emphasize informative phase regions, and a relative microphone positional encoding (rMPE), which encodes relative microphone coordinates in accordance with the nature of inter-channel time differences. Experiments on synthetic and real datasets demonstrate that AGG-RL achieves superior performance, particularly under unseen conditions. The results highlight the potential of representation learning with physics-informed design towards a universal solution for spatial acoustic scene understanding across diverse scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext da0928b1-1bce-447a-bd56-ac28fa4f58f9Builds on19
- Learning Audio-Visual Speech Representation by Masked Multimodal Cluster PredictionBowen Shi, Wei-Ning Hsu, Kushal Lakhotia, Abdelrahman MohamedICLR 2022 · 460 citations
- The Impact of Positional Encoding on Length Generalization in TransformersAmirhossein Kazemnejad, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Payel Das et al.NeurIPS 2023 · 444 citations
- Learning Neural Acoustic FieldsAndrew F. Luo, Yilun Du, Michael J. Tarr, Josh Tenenbaum et al.NeurIPS 2022 · 153 citations
- BinauralGrad: A Two-Stage Conditional Diffusion Probabilistic Model for Binaural Audio SynthesisYichong Leng, Zehua Chen, Junliang Guo, Haohe Liu et al.NeurIPS 2022 · 86 citations
- Neural Synthesis of Binaural Speech From Mono AudioAlexander Richard, Dejan Markovic, Israel D. Gebru, Steven Krenn et al.ICLR 2021 · 73 citations
Related papers
- PhaseCoder: Microphone Geometry-Agnostic Spatial Audio Understanding for Multimodal LLMsArtem Dementyev, Wazeer Zulfikar, Sinan Hersek, Pascal Getreuer et al.ICML 2026 · 4 citations
- Exploiting Transformation Invariance and Equivariance for Self-supervised Sound LocalisationJinxiang Liu, Chen Ju, Weidi Xie, Ya ZhangACM MM 2022 · 37 citations
- Self-Supervised Learning of Representations for Space Generates Multi-Modular Grid CellsRylan Schaeffer, Mikail Khona, Tzuhsuan Ma, Cristóbal Eyzaguirre et al.NeurIPS 2023 · 40 citations
- INRAS: Implicit Neural Representation for Audio ScenesKun Su, Mingfei Chen, Eli ShlizermanNeurIPS 2022 · 92 citations
- Radiance-Field Guided Pretraining: Scaling Localization Models with Unlabeled Wireless SignalsGuosheng Wang, Shen Wang, Lei YangUbiComp 2026
