Swarm Intelligence in Geo-Localization: A Multi-Agent Large Vision-Language Model Collaborative Framework
Xiao Han, Chen Zhu, Hengshu Zhu, Xiangyu Zhao
摘要
Visual geo-localization demands in-depth knowledge and advanced reasoning skills to associate images with precise real-world geographic locations. Existing image database retrieval methods are limited by the impracticality of storing sufficient visual records of global landmarks. Recently, Large Vision-Language Models (LVLMs) have demonstrated the capability of geo-localization through Visual Question Answering (VQA), enabling a solution that does not require external geo-tagged image records. However, the performance of a single LVLM is still limited by its intrinsic knowledge and reasoning capabilities. To address these challenges, we introduce smileGeo, a novel visual geo-localization framework that leverages multiple Internet-enabled LVLM agents operating within an agent-based architecture. By facilitating inter-agent communication, smileGeo integrates the inherent knowledge of these agents with additional retrieved information, enhancing the ability to effectively localize images. Furthermore, our framework incorporates a dynamic learning strategy that optimizes agent communication, reducing redundant interactions and enhancing overall system efficiency. To validate the effectiveness of the proposed framework, we conducted experiments on three different datasets, and the results show that our approach significantly outperforms current state-of-the-art methods. The source code is available at https://anonymous.4open.science/r/ViusalGeoLocalization-F8F5 . Relevance Statement: This paper focuses on invoking search, inferring the geo-locations of images through discussion and analysis of the retrieved information among multiple LVLMs, and responding in the form of natural language. The designed method in this paper can assist search and retrieval-augmented AI applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- GeoAgent: Learning to Geolocate Everywhere with Reinforced Geographic CharacteristicsModi Jin, Yiming Zhang, Boyuan Sun, Dingwen Zhang 等CVPR 2026 · 被引用 7 次
- GeoArena: Evaluating Open-World Geographic Reasoning in Large Vision-Language ModelsPengyue Jia, Yingyi Zhang, Xiangyu Zhao, Sharon LiACL 2026 · 被引用 3 次
- Multi-Agent Collaborative Reasoning with Tool-Augmented Evidence for Urban Region ProfilingXixuan Hao, Yutian Jiang, Jiabo Liu, Yihang Yang 等KDD 2026 · 被引用 2 次
- GeoBayes: Probabilistic Image Geo-Localization Inference via Sequential Bayesian UpdatingWeimin Shi, Xiang Li, Kaige Li, Junhao Fang 等AAAI 2026 · 被引用 2 次
- Do Vision-Language Models Respect Contextual Integrity in Location Disclosure?Ruixin Yang, Ethan Mendes, Arthur Wang, James Hays 等ICLR 2026 · 被引用 1 次
它引用的顶会 Paper17
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum 等ICML 2024 · 被引用 1,562 次
相关 Paper
- GRE Suite: Geo-localization Inference via Fine-Tuned Vision-Language Models and Enhanced Reasoning ChainsChun Wang, Xiaojun Ye, Xiaoran Pan, Zihao Pan 等NeurIPS 2025 · 被引用 18 次
- Recognition through Reasoning: Reinforcing Image Geo-localization with Large Vision-Language ModelsLing Li, Yao Zhou, Yuxuan Liang, Fugee Tsung 等NeurIPS 2025 · 被引用 30 次
- SpotAgent: Grounding Visual Geo-localization in Large Vision-Language Models through Agentic ReasoningFurong Jia, Ling Dai, Wenjin Deng, Fan Zhang 等KDD 2026 · 被引用 6 次
- Answering Complex Geographic Questions by Adaptive Reasoning with Visual Context and External Commonsense KnowledgeFan Li, Jianxing Yu, Jielong Tang, Wenqing Chen 等ACL 2025 · 被引用 3 次
- Large Language Models Know What is Key Visual Entity: An LLM-assisted Multimodal Retrieval for VQAPu Jian, Donglei Yu, Jiajun ZhangEMNLP 2024 · 被引用 5 次
