MMSciCode: Real-world Evaluation of Multilingual Multi-Discipline Scientific Research Coding
Xue Xia, Zheyuan Yang, Arman Cohan, Yilun Zhao
Abstract
We introduce MMSciCode, a comprehensive expert-level, multilingual multi-discipline benchmark for evaluating foundation models in scientific code generation. It includes 624 expert-annotated research coding problems spanning six core scientific disciplines. Compared to prior benchmarks, MMSciCode features three key advancements. First, it challenges models to integrate domain-specific knowledge with algorithmic reasoning to implement core functions from research papers. Second, each problem is meticulously annotated by domain experts through a rigorous paper-grounded process, with strict quality controls implemented to ensure dataset integrity and authenticity. Finally, each problem is equipped with comprehensive unit test suites and con-tainerized environments, enabling reproducible and diagnostic evaluation of both functional correctness and domain validity. We conduct an extensive evaluation of 23 state-of-the-art foundation models and 2 coding agents on MMSciCode. We identify substantial performance gaps between models and human experts, providing actionable insights for advancing expert-level scientific code generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 23e75142-02d7-4b5d-8a13-7cc4a5cb48a7Builds on7
- DS-1000: A Natural and Reliable Benchmark for Data Science Code GenerationYuhang Lai, Chengxi Li, Yiming Wang, Tianyi Zhang et al.ICML 2023 · 504 citations
- Paper2Code: Automating Code Generation from Scientific Papers in Machine LearningMinju Seo, Jinheon Baek, Seongyun Lee, Sung Ju HwangICLR 2026 · 86 citations
- RECODE-H: A Benchmark for Research Code Development with Interactive Human FeedbackChunyu Miao, Henry Peng Zou, Yangning Li, Yankai Chen et al.ICLR 2026 · 25 citations
- MLE-bench: Evaluating Machine Learning Agents on Machine Learning EngineeringJun Shern Chan, Neil Chowdhury, Oliver Jaffe, James Aung et al.ICLR 2025 · 9 citations
- LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling ResearchShuo Yan, Ruochen Li, Ziming Luo, Zimu Wang et al.EMNLP 2025
Related papers
- Omni-I2C: A Holistic Benchmark for High-Fidelity Image-to-Code GenerationJiawei Zhou, Chi Zhang, Xiang Feng, Qiming Zhang et al.ACL 2026 · 2 citations
- Top General Performance = Top Domain Performance? DomainCodeBench: A Multi-domain Code Generation BenchmarkDewu Zheng, Yanlin Wang, Ensheng Shi, Xilin Liu et al.ICSE 2026
- InteractScience: Programmatic and Visually-Grounded Evaluation of Interactive Scientific Demonstration Code GenerationQiaosheng Chen, Yang Liu, Lei Li, Kai Chen et al.ICML 2026 · 1 citation
- Uni-MMMU: A Massive Multi-discipline Multimodal Unified BenchmarkKai Zou, Ziqi Huang, Yuhao Dong, Shulin Tian et al.ACL 2026 · 19 citations
- DOMAINEVAL: An Auto-Constructed Benchmark for Multi-Domain Code GenerationQiming Zhu, Jialun Cao, Yaojie Lu, Hongyu Lin et al.AAAI 2025 · 25 citations
