InteractScience: Programmatic and Visually-Grounded Evaluation of Interactive Scientific Demonstration Code Generation
Qiaosheng Chen, Yang Liu, Lei Li, Kai Chen, Qipeng Guo, Gong Cheng, fei yuan
摘要
While Large Language Models (LLMs) hold promise for automating science and education, generating interactive scientific demonstrations demands a complex synthesis of deep domain knowledge and precise reactive coding. Current benchmarks fail to capture this synergy, largely bifurcating into static code generation or text-only reasoning. To address this, we introduce InteractScience, the first benchmark dedicated to evaluating the holistic creation of interactive scientific applications. We propose a novel hybrid framework that integrates programmatic functional testing for logic verification with visually-grounded qualitative assessment for rendering fidelity. Our evaluation of 30 leading models across five disciplines reveals critical gaps in grounding scientific reasoning within interactive interfaces. By standardizing this combined capability, InteractScience establishes a crucial foundation for reliable AI-driven tools in science and education.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- JanusCoder: Towards a Foundational Visual-Programmatic Interface for Code IntelligenceQiushi Sun, Jingyang Gong, Yang Liu, Qiaosheng Chen 等ICLR 2026 · 被引用 9 次
- RealChart2Code: Bridging the Gap in Real-World Chart-to-Code Generation via Multi-Task EvaluationJiajun Zhang, Yuying Li, Zhixun Li, Xingyu Guo 等ACL 2026
它引用的顶会 Paper12
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task SynthesisQiushi Sun, Kanzhi Cheng, Zichen Ding, Chuanyang Jin 等ACL 2025 · 被引用 114 次
- A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and RecommendationsMd. Tahmid Rahman Laskar, Sawsan Alqahtani, M. Saiful Bari, Mizanur Rahman 等EMNLP 2024 · 被引用 47 次
- WebCode2M: A Real-World Dataset for Code Generation from Webpage DesignsYi Gui, Zhen Li, Yao Wan, Yemin Shi 等WWW 2025 · 被引用 38 次
相关 Paper
- XDomainBench: Diagnosing Reasoning Collapse in High-Dimensional Scientific Knowledge CompositionZhiren Gong, Tiantong Wu, Jiaming Zhang, Fuyao Zhang 等ICML 2026
- MiniAppBench: Evaluating the Shift from Text to Interactive HTML Responses in LLM-Powered AssistantsZuhao zhang, Chengyue Yu, Yuante Li, Chenyi Zhuang 等ICML 2026 · 被引用 2 次
- LLM-SRBench: A New Benchmark for Scientific Equation Discovery with Large Language ModelsParshin Shojaee, Ngoc-Hieu Nguyen, Kazem Meidani, Amir Barati Farimani 等ICML 2025
- SCI-Verifier: Scientific Verifier with ThinkingShenghe Zheng, Chenyu Huang, Fangchen Yu, Junchi Yao 等ICLR 2026 · 被引用 5 次
- InteractBench: Benchmarking LLMs on Competitive Programming under Unrevealed InformationJiaze Li, Aocheng Shen, Bing Liu, Boyu Zhang 等ICML 2026
