SimulCost: A Cost-Aware Benchmark and Toolkit for Automating Physics Simulations with LLMs
Yadi Cao, Sicheng Lai, Jiahe Huang, Yang Zhang, Zach Lawrence, Rohan Bhakta, Izzy Thomas, Mingyun Cao, Chung-Hao Tsai, Zihao Zhou, Yidong Zhao, Hao Liu
Abstract
Evaluating LLM agents for scientific tasks has focused on token costs while ignoring tool-use costs like simulation time and experimental resources. As a result, metrics like pass@k become impractical under realistic budget constraints. To address this gap, we introduce SimulCost, the first benchmark targeting cost-sensitive parameter tuning in physics simulations. SimulCost compares LLM tuning cost-sensitive parameters against traditional scanning approach in both accuracy and computational cost, spanning 2,947 single-round (initial guess) and 1,931 multi-round (adjustment by trial-and-error) tasks across 13 simulators from fluid dynamics, solid mechanics, and plasma physics. Each simulator's cost is analytically defined and platform-independent. Frontier LLMs achieve 46--65% success rates in single-round mode, dropping to 35--55% under high accuracy requirements, rendering their initial guesses unreliable especially for high accuracy tasks. Multi-round mode improves rates to 72--81%, but LLMs are 1.5--2.5 slower than traditional scanning, making them uneconomical choices. We also investigate parameter group correlations for knowledge transfer potential, and the impact of in-context examples and reasoning effort, providing practical implications for deployment and fine-tuning. We open-source SimulCost as a static benchmark and extensible toolkit to facilitate research on improving cost-aware agentic designs for physics simulations, and for expanding new simulation environments. Code and data are available at https://github.com/Rose-STL-Lab/SimulCost-Bench.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c66b8a92-c988-4879-aee6-74385eccd948Builds on10
- DS-1000: A Natural and Reliable Benchmark for Data Science Code GenerationYuhang Lai, Chengxi Li, Yiming Wang, Tianyi Zhang et al.ICML 2023 · 504 citations
- Incremental potential contact: intersection-and inversion-free, large-deformation dynamicsMinchen Li, Zachary Ferguson, Teseo Schneider, Timothy R. Langlois et al.SIGGRAPH 2020 · 320 citations
- MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to UseYue Huang, Jiawen Shi, Yuan Li, Chenrui Fan et al.ICLR 2024 · 188 citations
- Design-Bench: Benchmarks for Data-Driven Offline Model-Based OptimizationBrandon Trabucco, Xinyang Geng, Aviral Kumar, Sergey LevineICML 2022 · 126 citations
- MLE-bench: Evaluating Machine Learning Agents on Machine Learning EngineeringJun Shern Chan, Neil Chowdhury, Oliver Jaffe, James Aung et al.ICLR 2025 · 9 citations
Related papers
- CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use AgentsJiayu Liu, Cheng Qian, Zhaochen Su, Qing Zong et al.ACL 2026 · 19 citations
- Crucible: Quantifying the Potential of Control Algorithms through LLM AgentsLianchen Jia, Chaoyang Li, Qian Houde, Tianchi Huang et al.NeurIPS 2025 · 1 citation
- SimWorld: An Open-ended Simulator for Agents in Physical and Social WorldsXiaokang Ye, Jiawei Ren, Yan Zhuang, Xuhong He et al.NeurIPS 2025 · 8 citations
- SimuHome: A Temporal- and Environment-Aware Benchmark for Smart Home LLM AgentsGyuhyeon Seo, Jungwoo Yang, Junseong Pyo, Nalim Kim et al.ICLR 2026 · 14 citations
- AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World ContextsKeyu Li, Junhao Shi, Yang Xiao, Mohan Jiang et al.ACL 2026 · 14 citations
