SEDRAS: Symbolically Evaluated Deep Research And Science
Fredrik Carlsson, Dan Ward, Joseph Ortiz, Fangyu Liu, Joakim Nivre
Abstract
As the reasoning capabilities of Large Language Models (LLMs) expand, evaluating true inductive generalization on entirely unseen data becomes increasingly challenging. To this end, we introduce a modular in-context learning evaluation framework, that is scalable and extendable across its separate modules. This is based upon the notion of synthetic scenarios with controllable complexity across three independent axes: 1) the logic of the underlying data distribution (UDD) 2) their projection into diverse representations, and 3) the interaction dynamic determining how the model accesses and explores the data. For these scenarios, the model is tasked to perform in-context scientific discovery and produce an interpretable theory in natural language that explains the observations. In a separate conversation, the model is then tasked to convert this generated theory into executable code, which can be programmatically compared against the underlying data distribution. Using this modular framework we produce an initial suite of 600 diverse scenarios that we use to evaluate and analyze various state-of-the-art LLMs. Although these experiments show that Gemini 3.0 Pro achieves the best overall score, each model performs the best at different tasks. For example: GPT 5.2 is the clear winner on pure symbolic data, Claude Opus 4.5 is the best at working with files, Gemini is the strongest model for the non-dynamic scenarios, and Grok 4.1 is the strongest model when UDD complexity scales. Furthermore, all models struggle with active exploration and are seemingly incapable of identifying informative data points, resulting in less efficient exploration than a random baseline.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on7
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou et al.ICLR 2024 · 1,197 citations
- AgentBench: Evaluating LLMs as AgentsXiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu et al.ICLR 2024 · 748 citations
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun et al.ICLR 2024 · 716 citations
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research AgentsMingxuan Du, Benfeng Xu, Chiwei Zhu, Licheng Zhang et al.ICLR 2026 · 250 citations
- NewtonBench: Benchmarking Generalizable Scientific Law Discovery in LLM AgentsTianshi Zheng, Kelvin Kiu Wai Tam, Newt Nguyen Kim Hue Nam, Baixuan Xu et al.ICLR 2026 · 25 citations
Related papers
- LLM-SRBench: A New Benchmark for Scientific Equation Discovery with Large Language ModelsParshin Shojaee, Ngoc-Hieu Nguyen, Kazem Meidani, Amir Barati Farimani et al.ICML 2025
- AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World EnvironmentsZhiheng Xi, Dingwen Yang, Jiaqi Liu, Jixuan Huang et al.ACL 2026
- Investigating Advanced Reasoning of Large Language Models via Black-Box Environment InteractionCongchi Yin, Tianyi Wu, Yankai Shu, Alex Gu et al.ICML 2026 · 1 citation
- SLR: Automated Synthesis for Scalable Logical ReasoningLukas Helff, Ahmad Omar, Felix Friedrich, Antonia Wüst et al.ACL 2026 · 6 citations
- ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research AgentsManasi Sharma, Chen Bo Calvin Zhang, Chaithanya Bandi, Clinton Wang et al.ICLR 2026 · 83 citations
