CityEQA: A Hierarchical LLM Agent on Embodied Question Answering Benchmark in City Space
Yong Zhao, Kai Xu, Zhengqiu Zhu, Yue Hu, Zhiheng Zheng, Yingfeng Chen, Yatai Ji, Chen Gao, Yong Li, Jincai Huang
摘要
Embodied Question Answering (EQA) has primarily focused on indoor environments, leaving the complexities of urban settings-spanning environment, action, and perception-largely unexplored. To bridge this gap, we introduce CityEQA, a new task where an embodied agent answers open-vocabulary questions through active exploration in dynamic city spaces. To support this task, we present CityEQA-EC, the first benchmark dataset featuring 1,412 human-annotated tasks across six categories, grounded in a realistic 3D urban simulator. Moreover, we propose Planner-Manager-Actor (PMA), a novel agent tailored for CityEQA. PMA enables long-horizon planning and hierarchical task execution: the Planner breaks down the question answering into sub-tasks, the Manager maintains an object-centric cognitive map for spatial reasoning during the process control, and the specialized Actors handle navigation, exploration, and collection sub-tasks. Experiments demonstrate that PMA achieves 60.7% of human-level answering accuracy, significantly outperforming competitive baselines. While promising, the performance gap compared to humans highlights the need for enhanced visual reasoning in CityEQA. This work paves the way for future advancements in urban spatial intelligence. Dataset and code are available at https://github.com/ tsinghua-fib-lab/CityEQA.git .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Theory of Space: Can Foundation Models Construct Spatial Beliefs through Active Exploration?Pingyue Zhang, Zihan Huang, Yue Wang, Jieyu Zhang 等ICLR 2026 · 被引用 23 次
- Towards Autonomous UAV Visual Object Search in City Space: Benchmark and Agentic MethodologyYatai Ji, Zhengqiu Zhu, Yong Zhao, Beidan Liu 等AAAI 2026 · 被引用 8 次
- CityNav: A Large-Scale Dataset for Real-World Aerial NavigationJungdae Lee, Taiki Miyanishi, Shuhei Kurita, Koya Sakamoto 等ICCV 2025 · 被引用 8 次
- USTBench: Benchmarking and Dissecting Spatiotemporal Reasoning Capabilities of LLMs as Urban AgentsSiqi Lai, Yansong Ning, Zirui Yuan, Zhixi Chen 等ICLR 2026 · 被引用 7 次
- PhysVLM-AVR: Active Visual Reasoning for Multimodal Large Language Models in Physical EnvironmentsWeijie Zhou, Xuantang Xiong, Yi Peng, Manli Tao 等NeurIPS 2025 · 被引用 4 次
它引用的顶会 Paper10
- SG-Nav: Online 3D Scene Graph Prompting for LLM-based Zero-shot Object NavigationHang Yin, Xiuwei Xu, Zhenyu Wu, Jie Zhou 等NeurIPS 2024 · 被引用 215 次
- Language Models Meet World Models: Embodied Experiences Enhance Language ModelsJiannan Xiang, Tianhua Tao, Yi Gu, Tianmin Shu 等NeurIPS 2023 · 被引用 180 次
- AerialVLN: Vision-and-Language Navigation for UAVsShubo Liu, Hongsheng Zhang, Yuankai Qi, Peng Wang 等ICCV 2023 · 被引用 132 次
- UrbanCLIP: Learning Text-enhanced Urban Region Profiling with Contrastive Language-Image Pretraining from the WebYibo Yan, Haomin Wen, Siru Zhong, Wei Chen 等WWW 2024 · 被引用 124 次
- EarthVQA: Towards Queryable Earth via Relational Reasoning-Based Remote Sensing Visual Question AnsweringJunjue Wang, Zhuo Zheng, Zihang Chen, Ailong Ma 等AAAI 2024 · 被引用 72 次
相关 Paper
- OpenEQA: Embodied Question Answering in the Era of Foundation ModelsArjun Majumdar, Anurag Ajay, Xiaohan Zhang, Pranav Putta 等CVPR 2024 · 被引用 44 次
- Beyond the Destination: A Novel Benchmark for Exploration-Aware Embodied Question AnsweringKaixuan Jiang, Yang Liu, Weixing Chen, Jingzhou Luo 等ICCV 2025 · 被引用 4 次
- Extending Embodied Question Answering from Perception to DecisionXicheng Gong, Qiwei Li, Peiran Xu, Yadong MuCVPR 2026 · 被引用 1 次
- 3D Question Answering for City Scene UnderstandingPenglei Sun, Yaoxian Song, Xiang Liu, Xiaofei Yang 等ACM MM 2024 · 被引用 6 次
- BridgeEQA: Virtual Embodied Agents for Real Bridge InspectionsSubin Varghese, Joshua Gao, Asad Ur Rahman, Vedhus HoskereCVPR 2026 · 被引用 3 次
