Earth-Agent: Unlocking the Full Landscape of Earth Observation with Agents
Peilin Feng, Zhutao Lv, Junyan Ye, Xiaolei Wang, Xinjie Huo, Jinhua Yu, Wanghan Xu, Wenlong Zhang, Lei Bai, Conghui He, Weijia Li
摘要
Earth observation (EO) is essential for understanding the evolving states of the Earth system. Although recent MLLMs have advanced EO research, they still lack the capability to tackle complex tasks that require multi-step reasoning and the use of domain-specific tools. Agent-based methods offer a promising direction, but current attempts remain in their infancy, confined to RGB perception, shallow reasoning, and lacking systematic evaluation protocols. To overcome these limitations, we introduce Earth-Agent, the first agentic framework that unifies RGB and spectral EO data within an MCP-based tool ecosystem, enabling cross-modal, multi-step, and quantitative spatiotemporal reasoning beyond pretrained MLLMs. Earth-Agent supports complex scientific tasks such as geophysical parameter retrieval and quantitative spatiotemporal analysis by dynamically invoking expert tools and models across modalities. To support comprehensive evaluation, we further propose Earth-Bench, a benchmark of 248 expert-curated tasks with 13,729 images, spanning spectrum, products and RGB modalities, and equipped with a dual-level evaluation protocol that assesses both reasoning trajectories and final outcomes. We conduct comprehensive experiments varying different LLM backbones, comparisons with general agent frameworks, and comparisons with MLLMs on remote sensing benchmarks, demonstrating both the effectiveness and potential of Earth-Agent. Earth-Agent establishes a new paradigm for EO analysis, moving the field toward scientifically grounded, next-generation applications of LLMs in Earth observation. More information about Earth-Agent can be found at https://github.com/opendatalab/Earth-Agent
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- TerraScope: Pixel-Grounded Visual Reasoning for Earth ObservationYan Shu, Bin Ren, Zhitong Xiong, Xiao Xiang Zhu 等CVPR 2026 · 被引用 9 次
- Multi-Agent Collaborative Reasoning with Tool-Augmented Evidence for Urban Region ProfilingXixuan Hao, Yutian Jiang, Jiabo Liu, Yihang Yang 等KDD 2026 · 被引用 2 次
- CausalGame: Benchmarking Causal Thinking of LLM Agents in GamesZhenhao Chen, Yongqiang Chen, Chenxi Liu, Junchi Yu 等ICML 2026 · 被引用 2 次
- RSMeM: Knowledge-Enhanced Memory Evolution for Remote Sensing Agents with Systematic EvaluationBingxian Wu, Yu Zhang, Zonghao Guo, Tang Liu 等ACL 2026
它引用的顶会 Paper18
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun 等ICLR 2024 · 被引用 716 次
- SkyScript: A Large and Semantically Diverse Vision-Language Dataset for Remote SensingZhecheng Wang, Rajanie Prabha, Tianyuan Huang, Jiajun Wu 等AAAI 2024 · 被引用 167 次
- AI-Researcher: Autonomous Scientific InnovationJiabin Tang, Lianghao Xia, Zhonghang Li, Chao HuangNeurIPS 2025 · 被引用 101 次
- Remote Sensing Vision-Language Foundation Models without Annotations via Ground Remote AlignmentUtkarsh Mall, Cheng Perng Phoo, Meilin Kelsey Liu, Carl Vondrick 等ICLR 2024 · 被引用 90 次
相关 Paper
- GeoMMBench and GeoMMAgent: Toward Expert-Level Multimodal Intelligence in Geoscience and Remote SensingAoran Xiao, Shihao Cheng, Yonghao Xu, Yexian Ren 等CVPR 2026 · 被引用 6 次
- MSEarth: A Multimodal Benchmark for Earth Science Phenomenon Discovery with MLLMsXiangyu Zhao, Wanghan Xu, Bo Liu, Yuhao Zhou 等ACL 2026 · 被引用 5 次
- Zephyrus: An Agentic Framework for Weather ScienceSumanth Varambally, Marshall Fisher, Jas Thakker, Yiwei Chen 等ICLR 2026 · 被引用 9 次
- DeepEyesV2: Toward Agentic Multimodal ModelJack Hong, Chenxiao Zhao, ChengLIn Zhu, Weiheng Lu 等ICLR 2026 · 被引用 109 次
- MMSearch-Plus: Benchmarking Provenance-Aware Search for Multimodal Browsing AgentsXijia Tao, Yihua Teng, Xinxing Su, Xinyu Fu 等ICLR 2026 · 被引用 37 次
