GeoBayes: Probabilistic Image Geo-Localization Inference via Sequential Bayesian Updating
Weimin Shi, Xiang Li, Kaige Li, Junhao Fang, Qiang Zhou, Qichuan Geng, Zhong Zhou
摘要
Image geo-localization aims to determine the geographic location of a query image. While Multimodal Large Language Models (MLLMs) show potential for this task due to their rich world knowledge and explainability, they often struggle with confirmation bias, i.e., committing prematurely to potentially incorrect guesses driven by visual clues with diverse geographic likelihoods. In this paper, we propose GeoBayes, a novel training-free framework that formulates geo-localization as a Maximum a Posteriori (MAP) estimation task over multiple geographic hypotheses and performs probabilistic reasoning via sequential Bayesian updating. GeoBayes regards each visual object and its associated geographic clues as probabilistic evidence, integrating them iteratively through a Hypothesize-Verify-Update loop. At each step, it evaluates how new evidence supports existing hypotheses and updates their posterior probabilities, gradually converging on the most probable location. This allows GeoBayes to explicitly quantify and fuse the varied geographic probabilities implied by diverse visual elements, reducing the risk of overcommitting to misleading clues. Furthermore, considering the natural hierarchy of geographic labels (e.g., country, city), GeoBayes introduces a state memory mechanism that stores hypotheses, inference context, and evidence scores across levels. This design enables the framework to propagate prior knowledge across levels of the geographic hierarchy and incorporate geographic structural constraints into the Bayesian update process, achieving a coarse-to-fine geo-localization. Experiments on IM2GPS3k and YFCC4K show that GeoBayes improves MLLM-based geo-localization accuracy without extra training. This demonstrates the effectiveness of probabilistic reasoning for robust and interpretable geo-localization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- SpatialRGPT: Grounded Spatial Reasoning in Vision-Language ModelsAn-Chieh Cheng, Hongxu Yin, Yang Fu, Qiushan Guo 等NeurIPS 2024 · 被引用 412 次
- GeoCLIP: Clip-Inspired Alignment between Locations and Images for Effective Worldwide Geo-localizationVicente Vivanco Cepeda, Gaurav Kumar Nayak, Mubarak ShahNeurIPS 2023 · 被引用 303 次
- Cross-view Geo-localization with Layer-to-Layer TransformerHongji Yang, Xiufan Lu, Yingying ZhuNeurIPS 2021 · 被引用 231 次
- TransGeo: Transformer Is All You Need for Cross-view Image Geo-localizationSijie Zhu, Mubarak Shah, Chen ChenCVPR 2022 · 被引用 189 次
- GeoLLM: Extracting Geospatial Knowledge from Large Language ModelsRohin Manvi, Samar Khanna, Gengchen Mai, Marshall Burke 等ICLR 2024 · 被引用 104 次
相关 Paper
- Vision-Language Reasoning for Geolocalization: A Reinforcement Learning ApproachBiao Wu, Meng Fang, Ling Chen, Ke Xu 等AAAI 2026 · 被引用 2 次
- Recognition through Reasoning: Reinforcing Image Geo-localization with Large Vision-Language ModelsLing Li, Yao Zhou, Yuxuan Liang, Fugee Tsung 等NeurIPS 2025 · 被引用 30 次
- GRE Suite: Geo-localization Inference via Fine-Tuned Vision-Language Models and Enhanced Reasoning ChainsChun Wang, Xiaojun Ye, Xiaoran Pan, Zihao Pan 等NeurIPS 2025 · 被引用 18 次
- Generate, Analyze, and Refine: Training-Free Sound Source Localization via MLLM Meta-ReasoningSubin Park, Jung Uk KimCVPR 2026
- GeoViS: Geospatially Rewarded Visual Search for Remote Sensing Visual GroundingPeirong Zhang, Yidan Zhang, Luxiao Xu, Jinliang Lin 等CVPR 2026 · 被引用 3 次
