IMAIA: Interactive Maps AI Assistant for Travel Planning and Geo-Spatial Intelligence
Jieren Deng, Zhizhang Hu, Ziyan He, Aleksandar Cvetkovic, Pak-Kiu Chung, Dragomir Yankov, Chiqun Zhang
摘要
Map applications are still largely point-and-click, making it difficult to ask map-centric questions or connect what a camera sees to the surrounding geospatial context with view-conditioned inputs. We introduce IMAIA, an interactive Maps AI Assistant that enables natural-language interaction with both vector (street) maps and satellite imagery, and augments camera inputs with geospatial intelligence to help users understand the world. IMAIA comprises two complementary components. Maps Plus treats the map as first-class context by parsing tiled vector/satellite views into a grid-aligned representation that a language model can query to resolve deictic references (e.g., ``the flower-shaped building next to the park in the top-right''). Places AI Smart Assistant (PAISA) performs camera-aware place understanding by fusing image--place embeddings with geospatial signals (location, heading, proximity) to ground a scene, surface salient attributes, and generate concise explanations. A lightweight multi-agent design keeps latency low and exposes interpretable intermediate decisions. Across map-centric QA and camera-to-place grounding tasks, IMAIA improves accuracy and responsiveness over strong baselines while remaining practical for user-facing deployments. By unifying language, maps, and geospatial cues, IMAIA moves beyond scripted tools toward conversational mapping that is both spatially grounded and broadly usable.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao 等NeurIPS 2024 · 被引用 2,305 次
- SpatialRGPT: Grounded Spatial Reasoning in Vision-Language ModelsAn-Chieh Cheng, Hongxu Yin, Yang Fu, Qiushan Guo 等NeurIPS 2024 · 被引用 412 次
- GeoLM: Empowering Language Models for Geospatially Grounded Language UnderstandingZekun Li, Wenxuan Zhou, Yao-Yi Chiang, Muhao ChenEMNLP 2023 · 被引用 25 次
- SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning CapabilitiesBoyuan Chen, Zhuo Xu, Sean Kirmani, Brian Ichter 等CVPR 2024
相关 Paper
- GeoVisA11y: An AI-based Geovisualization Question-Answering System for Screen-Reader UsersChu Li, Rock Yuren Pang, Arnavi Chheda-Kothary, Ather Sharif 等CHI 2026 · 被引用 1 次
- StreetViewAI: Making Street View Accessible Using Context-Aware Multimodal AIJon E. Froehlich, Alexander J. Fiannaca, Nimer Jaber, Victor Tsaran 等UIST 2025 · 被引用 5 次
- TraveLLaMA: A Multimodal Travel Assistant with Large-Scale Dataset and Structured ReasoningMeng Chu, Yukang Chen, Haokun Gui, Shaozuo Yu 等AAAI 2026
- EarthDial: Turning Multi-sensory Earth Observations to Interactive DialoguesSagar Soni, Akshay Dudhane, Hiyam Debary, Mustansar Fiaz 等CVPR 2025
- Textual Supervision Enhances Geospatial Representations in Vision-Language ModelsMarcelo Sartori Locatelli, Fernando Tonucci, Jea Kwon, Luiz Felipe Vecchietti 等ICML 2026
