GeoBayes: Probabilistic Image Geo-Localization Inference via Sequential Bayesian Updating
Weimin Shi, Xiang Li, Kaige Li, Junhao Fang, Qiang Zhou, Qichuan Geng, Zhong Zhou
Abstract
Image geo-localization aims to determine the geographic location of a query image. While Multimodal Large Language Models (MLLMs) show potential for this task due to their rich world knowledge and explainability, they often struggle with confirmation bias, i.e., committing prematurely to potentially incorrect guesses driven by visual clues with diverse geographic likelihoods. In this paper, we propose GeoBayes, a novel training-free framework that formulates geo-localization as a Maximum a Posteriori (MAP) estimation task over multiple geographic hypotheses and performs probabilistic reasoning via sequential Bayesian updating. GeoBayes regards each visual object and its associated geographic clues as probabilistic evidence, integrating them iteratively through a Hypothesize-Verify-Update loop. At each step, it evaluates how new evidence supports existing hypotheses and updates their posterior probabilities, gradually converging on the most probable location. This allows GeoBayes to explicitly quantify and fuse the varied geographic probabilities implied by diverse visual elements, reducing the risk of overcommitting to misleading clues. Furthermore, considering the natural hierarchy of geographic labels (e.g., country, city), GeoBayes introduces a state memory mechanism that stores hypotheses, inference context, and evidence scores across levels. This design enables the framework to propagate prior knowledge across levels of the geographic hierarchy and incorporate geographic structural constraints into the Bayesian update process, achieving a coarse-to-fine geo-localization. Experiments on IM2GPS3k and YFCC4K show that GeoBayes improves MLLM-based geo-localization accuracy without extra training. This demonstrates the effectiveness of probabilistic reasoning for robust and interpretable geo-localization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on21
- SpatialRGPT: Grounded Spatial Reasoning in Vision-Language ModelsAn-Chieh Cheng, Hongxu Yin, Yang Fu, Qiushan Guo et al.NeurIPS 2024 · 412 citations
- GeoCLIP: Clip-Inspired Alignment between Locations and Images for Effective Worldwide Geo-localizationVicente Vivanco Cepeda, Gaurav Kumar Nayak, Mubarak ShahNeurIPS 2023 · 303 citations
- Cross-view Geo-localization with Layer-to-Layer TransformerHongji Yang, Xiufan Lu, Yingying ZhuNeurIPS 2021 · 231 citations
- TransGeo: Transformer Is All You Need for Cross-view Image Geo-localizationSijie Zhu, Mubarak Shah, Chen ChenCVPR 2022 · 189 citations
- GeoLLM: Extracting Geospatial Knowledge from Large Language ModelsRohin Manvi, Samar Khanna, Gengchen Mai, Marshall Burke et al.ICLR 2024 · 104 citations
Related papers
- Vision-Language Reasoning for Geolocalization: A Reinforcement Learning ApproachBiao Wu, Meng Fang, Ling Chen, Ke Xu et al.AAAI 2026 · 2 citations
- Recognition through Reasoning: Reinforcing Image Geo-localization with Large Vision-Language ModelsLing Li, Yao Zhou, Yuxuan Liang, Fugee Tsung et al.NeurIPS 2025 · 30 citations
- GRE Suite: Geo-localization Inference via Fine-Tuned Vision-Language Models and Enhanced Reasoning ChainsChun Wang, Xiaojun Ye, Xiaoran Pan, Zihao Pan et al.NeurIPS 2025 · 18 citations
- Generate, Analyze, and Refine: Training-Free Sound Source Localization via MLLM Meta-ReasoningSubin Park, Jung Uk KimCVPR 2026
- GeoViS: Geospatially Rewarded Visual Search for Remote Sensing Visual GroundingPeirong Zhang, Yidan Zhang, Luxiao Xu, Jinliang Lin et al.CVPR 2026 · 3 citations
