Assessing Large Language Models on Climate Information
Jannis Bulian, Mike S. Schäfer, Afra Amini, Heidi Lam, Massimiliano Ciaramita, Ben Gaiarin, Michelle Chen Huebscher, Christian Buck, Niels Mede, Markus Leippold, Nadine Strauß
Abstract
As Large Language Models (LLMs) rise in popularity, it is necessary to assess their capability in critically relevant domains. We present a comprehensive evaluation framework, grounded in science communication research, to assess LLM responses to questions about climate change. Our framework emphasizes both presentational and epistemological adequacy, offering a fine-grained analysis of LLM generations spanning 8 dimensions and 30 issues. Our evaluation task is a real-world example of a growing number of challenging problems where AI can complement and lift human performance. We introduce a novel protocol for scalable oversight that relies on AI Assistance and raters with relevant education. We evaluate several recent LLMs on a set of diverse climate questions. Our results point to a significant gap between surface and epistemological qualities of LLMs in the realm of climate communication.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 270a9cf3-65a5-4821-9016-71afdfa8de99Cited by top-tier papers6
- Can Global XAI Methods Reveal Injected Behaviours in LLMs? SHAP vs Rule Extraction vs RuleSHAPFrancesco SovranoKDD 2026 · 5 citations
- MMClima: A Framework for Multimodal Climate Science Data and EvaluationMuhammad Umer Sheikh, Hassan Abid, Khawar shehzad, Ufaq Khan et al.ICML 2026 · 2 citations
- GCA Framework: A GCC Countries-Grounded Dataset and Agentic Pipeline for Climate Decision SupportMuhammad Umer Sheikh, Khawar Shehzad, Salman Khan, Fahad Shahbaz Khan et al.ACL 2026 · 1 citation
- Eating for a Sustainable Planet: Personalized Sustainable Diet Recommendation via Constraint-Aware Decision-Making ModelingYing Jin, Weiqing Min, Mingyu Huang, Shuqiang JiangICML 2026
- What can large language models do for sustainable food?Anna T. Thomas, Adam Yee, Andrew Mayne, Maya B. Mathur et al.ICML 2025
Builds on6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
- Large Dual Encoders Are Generalizable RetrieversJianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai et al.EMNLP 2022 · 145 citations
- A Critical Evaluation of Evaluations for Long-form Question AnsweringFangyuan Xu, Yixiao Song, Mohit Iyyer, Eunsol ChoiACL 2023 · 25 citations
Related papers
- ClimaQA: An Automated Evaluation Framework for Climate Question Answering ModelsVeeramakali Vignesh Manivannan, Yasaman Jafari, Srikar Eranky, Spencer Ho et al.ICLR 2025
- EducationQ: Evaluating LLMs' Teaching Capabilities Through Multi-Agent Dialogue FrameworkYao Shi, Rongkeng Liang, Yong XuACL 2025 · 18 citations
- Trustworthy Medical Question Answering: An Evaluation-Centric SurveyYinuo Wang, Baiyang Wang, Robert E. Mercer, Frank Rudzicz et al.EMNLP 2025 · 2 citations
- CriticEval: Evaluating Large-scale Language Model as CriticTian Lan, Wenwei Zhang, Chen Xu, Heyan Huang et al.NeurIPS 2024 · 26 citations
- Can LLMs replace Neil deGrasse Tyson? Evaluating the Reliability of LLMs as Science CommunicatorsPrasoon Bajpai, Niladri Chatterjee, Subhabrata Dutta, Tanmoy ChakrabortyEMNLP 2024
