MSEarth: A Multimodal Benchmark for Earth Science Phenomenon Discovery with MLLMs
Xiangyu Zhao, Wanghan Xu, Bo Liu, Yuhao Zhou, Fenghua Ling, Ben Fei, Xiaoyu Yue, Lei Bai, Wenlong Zhang, Xiao-Ming Wu
Abstract
The rapid advancement of multimodal large language models (MLLMs) offers new opportunities for complex scientific challenges, yet their application in earth science-especially at the graduate level-remains underexplored due to a lack of benchmarks reflecting the depth and complexity of geoscientific reasoning. Existing datasets often rely on synthetic data or simple figure-caption pairs, failing to capture the nuanced reasoning required for real-world applications. To address this, we introduce MSEarth, a multimodal scientific dataset and benchmark curated from high-quality, openaccess publications. Covering the five major spheres of Earth science-atmosphere, cryosphere, hydrosphere, lithosphere, and biosphere-MSEarth features over 289K figures with refined captions enriched by contextual discussions and reasoning from the original papers. The benchmark supports tasks such as scientific figure captioning, multiple choice questions, and open-ended reasoning, providing a scalable, high-fidelity resource for developing and evaluating MLLMs in scientific reasoning. MSEarth is publicly available to foster further research and innovation: https: //github.com/xiangyu-mm/MSEarth .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e5b91dd7-1bdb-4be1-86ca-3e26e1833a69Cited by top-tier papers2
- RadarQA: Multi-modal Quality Analysis of Weather Radar ForecastsXuming He, Zhiyuan You, Junchao Gong, Couhua Liu et al.NeurIPS 2025 · 14 citations
- UniM: A Unified Any-to-Any Interleaved Multimodal BenchmarkYanlin Li, Minghui Guo, Kaiwen Zhang, Shize Zhang et al.CVPR 2026 · 10 citations
Builds on9
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 · 2,727 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGIXiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang et al.CVPR 2024 · 213 citations
- Improving Automatic VQA Evaluation Using Large Language ModelsOscar Mañas, Benno Krojer, Aishwarya AgrawalAAAI 2024 · 59 citations
Related papers
- EarthSE: A Benchmark Evaluating Earth Scientific Exploration Capability for Large Language ModelsWanghan Xu, Xiangyu Zhao, Yuhao Zhou, Xiaoyu Yue et al.ICLR 2026 · 1 citation
- Earth-Agent: Unlocking the Full Landscape of Earth Observation with AgentsPeilin Feng, Zhutao Lv, Junyan Ye, Xiaolei Wang et al.ICLR 2026 · 49 citations
- MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language ModelsJiacheng Ruan, Dan Jiang, Xian Gao, Ting Liu et al.AAAI 2026 · 3 citations
- SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language ModelsXiaoxuan Wang, Ziniu Hu, Pan Lu, Yanqiao Zhu et al.ICML 2024 · 220 citations
- MMSI-Bench: A Benchmark for Multi-Image Spatial IntelligenceSihan Yang, Runsen Xu, Yiman Xie, Sizhe Yang et al.ICLR 2026 · 195 citations
