CURIE: Evaluating LLMs on Multitask Scientific Long-Context Understanding and Reasoning
Hao Cui, Zahra Shamsi, Gowoon Cheon, Xuejian Ma, Shutong Li, Maria Tikhanovskaya, Peter Christian Norgaard, Nayantara Mudur, Martyna Beata Plomecka, Paul Raccuglia, Yasaman Bahri, Victor V. Albert
Abstract
Scientific problem-solving involves synthesizing information while applying expert knowledge. We introduce CURIE, a scientific long-Context Understanding, Reasoning and Information Extraction benchmark to measure the potential of Large Language Models (LLMs) in scientific problem-solving and assisting scientists in realistic workflows. This benchmark introduces ten challenging tasks with a total of 580 problems and solution pairs curated by experts in six disciplinesmaterials science, condensed matter physics, quantum computing, geospatial analysis, biodiversity, and proteins -covering both experimental and theoretical workflows in science. We evaluate a range of closed and open LLMs on tasks in CURIE which requires domain expertise, comprehension of long in-context information, and multi-step reasoning. While Gemini Flash 2.0 and Claude-3 show consistent high comprehension across domains, the popular GPT-4o and command-R+ fail dramatically on protein sequencing tasks. With the best performance at 32% there is much room for improvement for all models. We hope that insights gained from CURIE can guide the future development of LLMs in sciences. Evaluation code and data links in: https://github.com/google/curie ⋆ equal technical contribution, ⋄ work done as a student researcher at Google
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Characterizing Deep Research: A Benchmark and Formal DefinitionAbhinav Java, Ashmit Khandelwal, Sukruta Prakash Midigeshi, Aaron Halfaker et al.ICLR 2026 · 30 citations
- Gated KalmaNet: A Fading Memory Layer through Test-time Ridge RegressionLiangzu Peng, Aditya Chattopadhyay, Luca Zancato, Elvis Nunez et al.CVPR 2026 · 10 citations
- CMT-Benchmark: A Benchmark for Condensed Matter Theory Built by Expert ResearchersHaining Pan, James V. Roggeveen, Erez Berg, Juan Alvarez et al.ICLR 2026 · 6 citations
- Efficient numeracy in language models through single-token number embeddingsLinus Kreitner, Paul Hager, Jonathan Mengedoht, Georgios Kaissis et al.ICML 2026 · 5 citations
- RCP-Merging: Merging Long Chain-of-Thought Models with Domain-Specific Models by Considering Reasoning Capability as PriorJunyao Yang, Jianwei Wang, Huiping Zhuang, Cen Chen et al.AAAI 2026 · 1 citation
Builds on8
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Focused Transformer: Contrastive Training for Context ScalingSzymon Tworkowski, Konrad Staniszewski, Mikolaj Pacek, Yuhuai Wu et al.NeurIPS 2023 · 190 citations
- Evaluating Open-Domain Question Answering in the Era of Large Language ModelsEhsan Kamalloo, Nouha Dziri, Charles L. A. Clarke, Davood RafieiACL 2023 · 96 citations
- LongBench: A Bilingual, Multitask Benchmark for Long Context UnderstandingYushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu et al.ACL 2024 · 94 citations
- QASA: Advanced Question Answering on Scientific ArticlesYoonjoo Lee, Kyungjae Lee, Sunghyun Park, Dasol Hwang et al.ICML 2023 · 76 citations
Related papers
- SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language ModelsXiaoxuan Wang, Ziniu Hu, Pan Lu, Yanqiao Zhu et al.ICML 2024 · 220 citations
- Demystifying Scientific Problem-Solving in LLMs by Probing Knowledge and ReasoningAlan Li, Yixin Liu, Arpan Sarkar, Doug Downey et al.ICML 2026 · 4 citations
- LLM-SRBench: A New Benchmark for Scientific Equation Discovery with Large Language ModelsParshin Shojaee, Ngoc-Hieu Nguyen, Kazem Meidani, Amir Barati Farimani et al.ICML 2025
- ExpertLongBench: Benchmarking Language Models on Expert-Level Long-Form Generation Tasks with Structured ChecklistsJie Ruan, Inderjeet Nair, Shuyang Cao, Amy Liu et al.ICLR 2026 · 25 citations
- Have LLMs Advanced Enough? A Challenging Problem Solving Benchmark For Large Language ModelsDaman Arora, Himanshu Gaurav Singh, MausamEMNLP 2023 · 36 citations
