SpotAgent: Grounding Visual Geo-localization in Large Vision-Language Models through Agentic Reasoning
Furong Jia, Ling Dai, Wenjin Deng, Fan Zhang, Chen Hu, Daxin Jiang, Yu Liu
Abstract
Large Vision-Language Models (LVLMs) have demonstrated strong reasoning capabilities in geo-localization, yet they often struggle in real-world scenarios where visual cues are sparse, long-tailed, and highly ambiguous. Previous approaches, bound by internal knowledge, often fail to provide verifiable results, yielding confident but ungrounded predictions when faced with confounded evidence. To address these challenges, we propose SpotAgent, a framework that formalizes geo-localization into an agentic reasoning process that leverages expert-level reasoning to synergize visual interpretation with tool-assisted verification. SpotAgent actively explores and verifies visual cues by leveraging external tools (e.g., web search, maps) through a ReAct diagram. We introduce a 3-stage post-training pipeline starting with a Supervised Fine-Tuning (SFT) stage for basic alignment, followed by an Agentic Cold Start phase utilizing high-quality trajectories synthesized via a Multi-Agent framework, aiming to instill tool-calling expertise. Subsequently, the model's reasoning capabilities are refined through Reinforcement Learning. We propose a Spatially-Aware Dynamic Filtering strategy to enhance the efficiency of the RL stage by prioritizing learnable samples based on spatial difficulty. Extensive experiments on standard benchmarks demonstrate that SpotAgent achieves state-of-the-art performance, effectively mitigating hallucinations while delivering precise and verifiable geo-localization. Our code is available at: https://jiafr1802.github.io/SpotAgent-Paper/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6aa7dcdd-13c5-469d-9064-44be1cb3ceb6Builds on25
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- ViperGPT: Visual Inference via Python Execution for ReasoningDídac Surís, Sachit Menon, Carl VondrickICCV 2023 · 732 citations
- GeoCLIP: Clip-Inspired Alignment between Locations and Images for Effective Worldwide Geo-localizationVicente Vivanco Cepeda, Gaurav Kumar Nayak, Mubarak ShahNeurIPS 2023 · 303 citations
- Cross-view Geo-localization with Layer-to-Layer TransformerHongji Yang, Xiufan Lu, Yingying ZhuNeurIPS 2021 · 231 citations
Related papers
- Connecting the Dots: Training-Free Visual Grounding via Agentic ReasoningLiqin Luo, Guangyao Chen, Xiawu Zheng, Yongxing Dai et al.AAAI 2026
- Swarm Intelligence in Geo-Localization: A Multi-Agent Large Vision-Language Model Collaborative FrameworkXiao Han, Chen Zhu, Hengshu Zhu, Xiangyu ZhaoKDD 2025 · 1 citation
- RSAgent: Learning to Reason and Act via Multi-Turn Tool Invocations for Text-Guided SegmentationXingqi He, Yujie Zhang, Shuyong Gao, Wenjie Li et al.ICML 2026 · 3 citations
- GRE Suite: Geo-localization Inference via Fine-Tuned Vision-Language Models and Enhanced Reasoning ChainsChun Wang, Xiaojun Ye, Xiaoran Pan, Zihao Pan et al.NeurIPS 2025 · 18 citations
- FactGuard: Agentic Video Misinformation Detection via Reinforcement LearningZehao Li, Hongwei Yu, Hao Jiang, Qiang Sheng et al.ICML 2026 · 3 citations
