Can LLMs replace Neil deGrasse Tyson? Evaluating the Reliability of LLMs as Science Communicators
Prasoon Bajpai, Niladri Chatterjee, Subhabrata Dutta, Tanmoy Chakraborty
Abstract
Large Language Models (LLMs) and AI assistants driven by these models are experiencing exponential growth in usage among both expert and amateur users. In this work, we focus on evaluating the reliability of current LLMs as science communicators. Unlike existing benchmarks, our approach emphasizes assessing these models on scientific questionanswering tasks that require a nuanced understanding and awareness of answerability. We introduce a novel dataset, SCiPS-QA, comprising 742 Yes/No queries embedded in complex scientific concepts, along with a benchmarking suite that evaluates LLMs for correctness and consistency across various criteria. We benchmark three proprietary LLMs from the OpenAI GPT family and 13 open-access LLMs from the Meta Llama-2, Llama-3, and Mistral families. While most open-access models significantly underperform compared to GPT-4 Turbo, our experiments identify Llama-3-70B as a strong competitor, often surpassing GPT-4 Turbo in various evaluation aspects. We also find that even the GPT models exhibit a general incompetence in reliably verifying LLM responses. Moreover, we observe an alarming trend where human evaluators are deceived by incorrect responses from GPT-4 Turbo.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- PaperTrail: A Claim-Evidence Interface for Grounding Provenance in LLM-based Scholarly Q&AAnna Martin-Boyle, Cara A. C. Leckey, Martha Brown, Harmanpreet KaurCHI 2026 · 3 citations
- An Expert Schema for Evaluating Large Language Model Errors in Scholarly Question-Answering SystemsAnna Martin-Boyle, William Humphreys, Martha Brown, Cara A. C. Leckey et al.CHI 2026 · 1 citation
Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 · 2,727 citations
- Large Language Models Cannot Self-Correct Reasoning YetJie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng et al.ICLR 2024 · 858 citations
Related papers
- SciAgent: Tool-augmented Language Models for Scientific ReasoningYubo Ma, Zhibin Gou, Junheng Hao, Ruochen Xu et al.EMNLP 2024 · 13 citations
- SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language ModelsXiaoxuan Wang, Ziniu Hu, Pan Lu, Yanqiao Zhu et al.ICML 2024 · 220 citations
- FactBench: A Dynamic Benchmark for In-the-Wild Language Model Factuality EvaluationFarima Fatahi Bayat, Lechen Zhang, Sheza Munir, Lu WangACL 2025
- SciEx: Benchmarking Large Language Models on Scientific Exams with Human Expert Grading and Automatic GradingTu Anh Dinh, Carlos Mullov, Leonard Bärmann, Zhaolin Li et al.EMNLP 2024 · 10 citations
- SciEval: A Multi-Level Large Language Model Evaluation Benchmark for Scientific ResearchLiangtai Sun, Yang Han, Zihan Zhao, Da Ma et al.AAAI 2024 · 150 citations
