PSM: Learning Probabilistic Embeddings for Multi-scale Zero-Shot Soundscape Mapping
Subash Khanal, Eric Xing, Srikumar Sastry, Aayush Dhakal, Zhexiao Xiong, Adeel Ahmad, Nathan Jacobs
Abstract
A soundscape is defined by the acoustic environment a person perceives at a location. In this work, we propose a framework for mapping soundscapes across the Earth. Since soundscapes involve sound distributions that span varying spatial scales, we represent locations with multi-scale satellite imagery and learn a joint representation among this imagery, audio, and text. To capture the inherent uncertainty in the soundscape of a location, we design the representation space to be probabilistic. We also fuse ubiquitous metadata (including geolocation, time, and data source) to enable learning of spatially and temporally dynamic representations of soundscapes. We demonstrate the utility of our framework by creating large-scale soundscape maps integrating both audio and text with temporal control. To facilitate future research on this task, we also introduce a large-scale dataset, GeoSound, containing over 300k geotagged audio samples paired with both low- and high-resolution satellite imagery. We demonstrate that our method outperforms the existing state-of-the-art on both GeoSound and the existing SoundingEarth dataset. Our dataset and code is available at https://github.com/mvrl/PSM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 89ab6d7c-2676-4ff8-a12e-298ebb0c9159Cited by top-tier papers3
- Bioacoustic Geolocation: Species Sounds as Geographic SignalsMustafa Chasmai, Wuao Liu, Subhransu Maji, Grant HornICML 2026 · 5 citations
- ProM3E: Probabilistic Masked MultiModal Embedding Model for EcologySrikumar Sastry, Subash Khanal, Aayush Dhakal, Jiayu Lin et al.CVPR 2026 · 1 citation
- RANGE: Retrieval Augmented Neural Fields for Multi-Resolution Geo-EmbeddingsAayush Dhakal, Srikumar Sastry, Subash Khanal, Adeel Ahmad et al.CVPR 2025
Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- SatMAE: Pre-training Transformers for Temporal and Multi-Spectral Satellite ImageryYezhen Cong, Samar Khanna, Chenlin Meng, Patrick Liu et al.NeurIPS 2022 · 707 citations
- Scale-MAE: A Scale-Aware Masked Autoencoder for Multiscale Geospatial Representation LearningColorado J. Reed, Ritwik Gupta, Shufan Li, Sarah Brockman et al.ICCV 2023 · 373 citations
- X-CLIP: End-to-End Multi-grained Contrastive Learning for Video-Text RetrievalYiwei Ma, Guohai Xu, Xiaoshuai Sun, Ming Yan et al.ACM MM 2022 · 314 citations
- Learning with Noisy Correspondence for Cross-modal MatchingZhenyu Huang, Guocheng Niu, Xiao Liu, Wenbiao Ding et al.NeurIPS 2021 · 215 citations
Related papers
- SounDiT: Geo-Contextual Soundscape-to-Landscape GenerationJunbo Wang, Haofeng Tan, Bowen Liao, Albert Jiang et al.CVPR 2026 · 3 citations
- GeoSURGE: Geo-localization using Semantic Fusion with Hierarchy of Geographic EmbeddingsAngel Daruna, Nicholas Meegan, Han-Pang Chiu, Supun Samarasekera et al.CVPR 2026 · 2 citations
- Learning Spatially-Aware Language and Audio EmbeddingsBhavika Devnani, Skyler Seto, Zakaria Aldeneh, Alessandro Toso et al.NeurIPS 2024 · 31 citations
- GeoRanker: Distance-Aware Ranking for Worldwide Image GeolocalizationPengyue Jia, Seongheon Park, Song Gao, Xiangyu Zhao et al.NeurIPS 2025 · 22 citations
- SOUPLE: Enhancing Audio-Visual Localization and Segmentation with Learnable Prompt ContextsKhanh Binh Nguyen, Chae Jung ParkCVPR 2026
