InnoGym: Benchmarking the Innovation Potential of AI Agents
Jintian Zhang, Kewei Xu, Jingsheng Zheng, Zhuoyun Yu, Yuqi Zhu, Yujie Luo, Lanning Wei, Shuofei Qiao, Lun Du, Da Zheng, Shumin Deng, Huajun Chen, Ningyu Zhang
Abstract
LLMs and Agents have achieved impressive progress in code generation, mathematical reasoning, and scientific discovery. However, existing benchmarks primarily measure correctness, overlooking the diversity of methods behind solutions. True innovation depends not only on producing correct answers but also on the originality of the approach. We present InnoGym, the first benchmark and framework designed to systematically evaluate the innovation potential of AI agents. InnoGym introduces two complementary metrics: performance gain, which measures improvement over the best-known solutions, and novelty, which captures methodological differences from prior approaches. The benchmark includes 18 carefully curated tasks from real-world engineering and scientific domains, each standardized through resource filtering, evaluator validation, and solution collection. In addition, we provide iGym, a unified execution environment for reproducible and long-horizon evaluations. Extensive experiments show that while some agents produce novel approaches, their lack of robustness limits performance gains. These results highlight a key gap between creativity and effectiveness, underscoring the need for benchmarks that evaluate both. Building on this framework, we present InnoGym, which consists of two complementary components: iBench and iGym. iBench is the first benchmark specifically designed to evaluate the innovation potential of AI agents. It includes 18 carefully curated Improvable Tasks, selected from † Corresponding author.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4ab6e9b0-9651-4102-9ead-4886c47cce47Builds on14
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- Neural Sparse Voxel FieldsLingjie Liu, Jiatao Gu, Kyaw Zaw Lin, Tat-Seng Chua et al.NeurIPS 2020 · 1,535 citations
Related papers
- InnovatorBench: Evaluating Agents' Ability to Conduct Innovative AI ResearchYunze Wu, Dayuan Fu, Weiye Si, Zhen Huang et al.ICLR 2026 · 9 citations
- Assessing the Creativity of LLMs in Proposing Novel Solutions to Mathematical ProblemsJunyi Ye, Jingyi Gu, Xinyun Zhao, Wenpeng Yin et al.AAAI 2025 · 2 citations
- Automated Creativity Evaluation of Language Models Across Open-Ended TasksTan Min Sen, Zachary Choy Kit Chun, Syed Ali Redha Alsagoff, Nadya Yuki Wangsajaya et al.ACL 2026
- DevOps-Gym: Benchmarking AI Agents in Software DevOps CycleYuheng Tang, Kaijie Zhu, Bonan Ruan, Chuqi Zhang et al.ICLR 2026 · 10 citations
- HeuriGym: An Agentic Benchmark for LLM-Crafted Heuristics in Combinatorial OptimizationHongzheng Chen, Yingheng Wang, Yaohui Cai, Hins Hu et al.ICLR 2026 · 26 citations
