ScienceWorld: Is your Agent Smarter than a 5th Grader?
Ruoyao Wang, Peter A. Jansen, Marc-Alexandre Côté, Prithviraj Ammanabrolu
Abstract
We present ScienceWorld, a benchmark to test agents' scientific reasoning abilities in a new interactive text environment at the level of a standard elementary school science curriculum. Despite the transformer-based progress seen in question-answering and scientific text processing, we find that current models cannot reason about or explain learned science concepts in novel contexts. For instance, models can easily answer what the conductivity of a known material is but struggle when asked how they would conduct an experiment in a grounded environment to find the conductivity of an unknown material. This begs the question of whether current models are simply retrieving answers by way of seeing a large number of similar examples or if they have learned to reason about concepts in a reusable manner. We hypothesize that agents need to be grounded in interactive environments to achieve such reasoning capabilities. Our experiments provide empirical evidence supporting this hypothesis -- showing that a 1.5 million parameter agent trained interactively for 100k steps outperforms a 11 billion parameter model statically trained for scientific question-answering and reasoning from millions of expert demonstrations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers60
- Grounding Large Language Models in Interactive Environments with Online Reinforcement LearningThomas Carta, Clément Romac, Thomas Wolf, Sylvain Lamprier et al.ICML 2023 · 258 citations
- SwiftSage: A Generative Agent with Fast and Slow Thinking for Complex Interactive TasksBill Yuchen Lin, Yicheng Fu, Karina Yang, Faeze Brahman et al.NeurIPS 2023 · 244 citations
- Evolving AgentsLeonardo RanaldiACL 2026 · 227 citations
- Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement LearningSimon Zhai, Hao Bai, Zipeng Lin, Jiayi Pan et al.NeurIPS 2024 · 214 citations
- MemGen: Weaving Generative Latent Memory for Self-Evolving AgentsGuibin Zhang, Muxin Fu, Shuicheng YanICLR 2026 · 102 citations
Builds on14
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied AgentsWenlong Huang, Pieter Abbeel, Deepak Pathak, Igor MordatchICML 2022 · 1,539 citations
- ALFWorld: Aligning Text and Embodied Environments for Interactive LearningMohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk et al.ICLR 2021 · 819 citations
- QASC: A Dataset for Question Answering via Sentence CompositionTushar Khot, Peter Clark, Michal Guerquin, Peter Jansen et al.AAAI 2020 · 387 citations
- Interactive Fiction Games: A Colossal AdventureMatthew J. Hausknecht, Prithviraj Ammanabrolu, Marc-Alexandre Côté, Xingdi YuanAAAI 2020 · 242 citations
Related papers
- Exploring the Capacity of Pretrained Language Models for Reasoning about Actions and ChangeWeinan He, Canming Huang, Zhanhao Xiao, Yongmei LiuACL 2023
- SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language ModelsXiaoxuan Wang, Ziniu Hu, Pan Lu, Yanqiao Zhu et al.ICML 2024 · 220 citations
- InteractScience: Programmatic and Visually-Grounded Evaluation of Interactive Scientific Demonstration Code GenerationQiaosheng Chen, Yang Liu, Lei Li, Kai Chen et al.ICML 2026 · 1 citation
- Benchmarking World-Model Learning with Environment-Level QueriesArchana Warrier, Dat Nguyen, Michelangelo Naim, Moksh Jain et al.ICML 2026 · 3 citations
- SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal ModelsAndong Deng, Taojiannan Yang, Shoubin Yu, Lincoln Spencer et al.ICML 2026 · 7 citations
