InteractScience: Programmatic and Visually-Grounded Evaluation of Interactive Scientific Demonstration Code Generation
Qiaosheng Chen, Yang Liu, Lei Li, Kai Chen, Qipeng Guo, Gong Cheng, fei yuan
Abstract
While Large Language Models (LLMs) hold promise for automating science and education, generating interactive scientific demonstrations demands a complex synthesis of deep domain knowledge and precise reactive coding. Current benchmarks fail to capture this synergy, largely bifurcating into static code generation or text-only reasoning. To address this, we introduce InteractScience, the first benchmark dedicated to evaluating the holistic creation of interactive scientific applications. We propose a novel hybrid framework that integrates programmatic functional testing for logic verification with visually-grounded qualitative assessment for rendering fidelity. Our evaluation of 30 leading models across five disciplines reveals critical gaps in grounding scientific reasoning within interactive interfaces. By standardizing this combined capability, InteractScience establishes a crucial foundation for reliable AI-driven tools in science and education.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 085d217a-41fe-4060-8b70-698f77a83ea6Cited by top-tier papers2
- JanusCoder: Towards a Foundational Visual-Programmatic Interface for Code IntelligenceQiushi Sun, Jingyang Gong, Yang Liu, Qiaosheng Chen et al.ICLR 2026 · 9 citations
- RealChart2Code: Bridging the Gap in Real-World Chart-to-Code Generation via Multi-Task EvaluationJiajun Zhang, Yuying Li, Zhixun Li, Xingyu Guo et al.ACL 2026
Builds on12
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task SynthesisQiushi Sun, Kanzhi Cheng, Zichen Ding, Chuanyang Jin et al.ACL 2025 · 114 citations
- A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and RecommendationsMd. Tahmid Rahman Laskar, Sawsan Alqahtani, M. Saiful Bari, Mizanur Rahman et al.EMNLP 2024 · 47 citations
- WebCode2M: A Real-World Dataset for Code Generation from Webpage DesignsYi Gui, Zhen Li, Yao Wan, Yemin Shi et al.WWW 2025 · 38 citations
Related papers
- XDomainBench: Diagnosing Reasoning Collapse in High-Dimensional Scientific Knowledge CompositionZhiren Gong, Tiantong Wu, Jiaming Zhang, Fuyao Zhang et al.ICML 2026
- MiniAppBench: Evaluating the Shift from Text to Interactive HTML Responses in LLM-Powered AssistantsZuhao zhang, Chengyue Yu, Yuante Li, Chenyi Zhuang et al.ICML 2026 · 2 citations
- LLM-SRBench: A New Benchmark for Scientific Equation Discovery with Large Language ModelsParshin Shojaee, Ngoc-Hieu Nguyen, Kazem Meidani, Amir Barati Farimani et al.ICML 2025
- SCI-Verifier: Scientific Verifier with ThinkingShenghe Zheng, Chenyu Huang, Fangchen Yu, Junchi Yao et al.ICLR 2026 · 5 citations
- InteractBench: Benchmarking LLMs on Competitive Programming under Unrevealed InformationJiaze Li, Aocheng Shen, Bing Liu, Boyu Zhang et al.ICML 2026
