IMAIA: Interactive Maps AI Assistant for Travel Planning and Geo-Spatial Intelligence
Jieren Deng, Zhizhang Hu, Ziyan He, Aleksandar Cvetkovic, Pak-Kiu Chung, Dragomir Yankov, Chiqun Zhang
Abstract
Map applications are still largely point-and-click, making it difficult to ask map-centric questions or connect what a camera sees to the surrounding geospatial context with view-conditioned inputs. We introduce IMAIA, an interactive Maps AI Assistant that enables natural-language interaction with both vector (street) maps and satellite imagery, and augments camera inputs with geospatial intelligence to help users understand the world. IMAIA comprises two complementary components. Maps Plus treats the map as first-class context by parsing tiled vector/satellite views into a grid-aligned representation that a language model can query to resolve deictic references (e.g., ``the flower-shaped building next to the park in the top-right''). Places AI Smart Assistant (PAISA) performs camera-aware place understanding by fusing image--place embeddings with geospatial signals (location, heading, proximity) to ground a scene, surface salient attributes, and generate concise explanations. A lightweight multi-agent design keeps latency low and exposes interpretable intermediate decisions. Across map-centric QA and camera-to-place grounding tasks, IMAIA improves accuracy and responsiveness over strong baselines while remaining practical for user-facing deployments. By unifying language, maps, and geospatial cues, IMAIA moves beyond scripted tools toward conversational mapping that is both spatially grounded and broadly usable.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on6
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
- SpatialRGPT: Grounded Spatial Reasoning in Vision-Language ModelsAn-Chieh Cheng, Hongxu Yin, Yang Fu, Qiushan Guo et al.NeurIPS 2024 · 412 citations
- GeoLM: Empowering Language Models for Geospatially Grounded Language UnderstandingZekun Li, Wenxuan Zhou, Yao-Yi Chiang, Muhao ChenEMNLP 2023 · 25 citations
- SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning CapabilitiesBoyuan Chen, Zhuo Xu, Sean Kirmani, Brian Ichter et al.CVPR 2024
Related papers
- GeoVisA11y: An AI-based Geovisualization Question-Answering System for Screen-Reader UsersChu Li, Rock Yuren Pang, Arnavi Chheda-Kothary, Ather Sharif et al.CHI 2026 · 1 citation
- StreetViewAI: Making Street View Accessible Using Context-Aware Multimodal AIJon E. Froehlich, Alexander J. Fiannaca, Nimer Jaber, Victor Tsaran et al.UIST 2025 · 5 citations
- TraveLLaMA: A Multimodal Travel Assistant with Large-Scale Dataset and Structured ReasoningMeng Chu, Yukang Chen, Haokun Gui, Shaozuo Yu et al.AAAI 2026
- EarthDial: Turning Multi-sensory Earth Observations to Interactive DialoguesSagar Soni, Akshay Dudhane, Hiyam Debary, Mustansar Fiaz et al.CVPR 2025
- Textual Supervision Enhances Geospatial Representations in Vision-Language ModelsMarcelo Sartori Locatelli, Fernando Tonucci, Jea Kwon, Luiz Felipe Vecchietti et al.ICML 2026
