RANGE: Retrieval Augmented Neural Fields for Multi-Resolution Geo-Embeddings
Aayush Dhakal, Srikumar Sastry, Subash Khanal, Adeel Ahmad, Eric Xing, Nathan Jacobs
摘要
The choice of representation for geographic location significantly impacts the accuracy of models for a broad range of geospatial tasks, including fine-grained species classification, population density estimation, and biome classification. Recent works like SatCLIP and GeoCLIP learn such representations by contrastively aligning geolocation with co-located images. While these methods work exceptionally well, in this paper, we posit that the current training strategies fail to fully capture the important visual features. We provide an information theoretic perspective on why the resulting embeddings from these methods discard crucial visual information that is important for many downstream tasks. To solve this problem, we propose a novel retrievalaugmented strategy called RANGE. We build our method on the intuition that the visual features of a location can be estimated by combining the visual features from multiple similar-looking locations. We evaluate our method across a wide variety of tasks. Our results show that RANGE outperforms the existing state-of-the-art models with significant margins in most tasks. We show gains of up to 13.1% on classification tasks and 0.145 R 2 on regression tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Measuring the Intrinsic Dimension of Earth RepresentationsArjun Rao, Marc Rußwurm, Konstantin Klemmer, Esther RolfICLR 2026 · 被引用 13 次
- Localized, High-resolution Geographic Representations with Slepian FunctionsArjun Rao, Ruth Crasto, Tessa Ooms, David Rolnick 等ICML 2026 · 被引用 2 次
- UrbanFusion: Stochastic Multimodal Fusion for Contrastive Learning of Robust Spatial RepresentationsDominik J. Mühlematter, Lin Che, Ye Hong, Martin Raubal 等ICML 2026
- Beyond What's Shared: Recovering Lost Unique Information from Intermediate Layers to Boost Multimodal Geo-Foundation ModelsJangHyeon Lee, Philipe Ambrozio Dias, Yao-Yi Chiang, Dalton LungaCVPR 2026
它引用的顶会 Paper23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai 等ICML 2022 · 被引用 1,629 次
相关 Paper
- SatCLIP: Global, General-Purpose Location Embeddings with Satellite ImageryKonstantin Klemmer, Esther Rolf, Caleb Robinson, Lester Mackey 等AAAI 2025 · 被引用 173 次
- G3: An Effective and Adaptive Framework for Worldwide Geolocalization Using Large Multi-Modality ModelsPengyue Jia, Yiding Liu, Xiaopeng Li, Xiangyu Zhao 等NeurIPS 2024 · 被引用 60 次
- GeoCLIP: Clip-Inspired Alignment between Locations and Images for Effective Worldwide Geo-localizationVicente Vivanco Cepeda, Gaurav Kumar Nayak, Mubarak ShahNeurIPS 2023 · 被引用 303 次
- Scaling Image Geo-Localization to Continent LevelPhilipp Lindenberger, Paul-Edouard Sarlin, Jan Hosang, Marc Pollefeys 等NeurIPS 2025 · 被引用 11 次
- Viewpoint Invariant Dense Matching for Visual GeolocalizationGabriele Moreno Berton, Carlo Masone, Valerio Paolicelli, Barbara CaputoICCV 2021 · 被引用 48 次
