SimulCost: A Cost-Aware Benchmark and Toolkit for Automating Physics Simulations with LLMs
Yadi Cao, Sicheng Lai, Jiahe Huang, Yang Zhang, Zach Lawrence, Rohan Bhakta, Izzy Thomas, Mingyun Cao, Chung-Hao Tsai, Zihao Zhou, Yidong Zhao, Hao Liu
摘要
Evaluating LLM agents for scientific tasks has focused on token costs while ignoring tool-use costs like simulation time and experimental resources. As a result, metrics like pass@k become impractical under realistic budget constraints. To address this gap, we introduce SimulCost, the first benchmark targeting cost-sensitive parameter tuning in physics simulations. SimulCost compares LLM tuning cost-sensitive parameters against traditional scanning approach in both accuracy and computational cost, spanning 2,947 single-round (initial guess) and 1,931 multi-round (adjustment by trial-and-error) tasks across 13 simulators from fluid dynamics, solid mechanics, and plasma physics. Each simulator's cost is analytically defined and platform-independent. Frontier LLMs achieve 46--65% success rates in single-round mode, dropping to 35--55% under high accuracy requirements, rendering their initial guesses unreliable especially for high accuracy tasks. Multi-round mode improves rates to 72--81%, but LLMs are 1.5--2.5 slower than traditional scanning, making them uneconomical choices. We also investigate parameter group correlations for knowledge transfer potential, and the impact of in-context examples and reasoning effort, providing practical implications for deployment and fine-tuning. We open-source SimulCost as a static benchmark and extensible toolkit to facilitate research on improving cost-aware agentic designs for physics simulations, and for expanding new simulation environments. Code and data are available at https://github.com/Rose-STL-Lab/SimulCost-Bench.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- DS-1000: A Natural and Reliable Benchmark for Data Science Code GenerationYuhang Lai, Chengxi Li, Yiming Wang, Tianyi Zhang 等ICML 2023 · 被引用 504 次
- Incremental potential contact: intersection-and inversion-free, large-deformation dynamicsMinchen Li, Zachary Ferguson, Teseo Schneider, Timothy R. Langlois 等SIGGRAPH 2020 · 被引用 320 次
- MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to UseYue Huang, Jiawen Shi, Yuan Li, Chenrui Fan 等ICLR 2024 · 被引用 188 次
- Design-Bench: Benchmarks for Data-Driven Offline Model-Based OptimizationBrandon Trabucco, Xinyang Geng, Aviral Kumar, Sergey LevineICML 2022 · 被引用 126 次
- MLE-bench: Evaluating Machine Learning Agents on Machine Learning EngineeringJun Shern Chan, Neil Chowdhury, Oliver Jaffe, James Aung 等ICLR 2025 · 被引用 9 次
相关 Paper
- CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use AgentsJiayu Liu, Cheng Qian, Zhaochen Su, Qing Zong 等ACL 2026 · 被引用 19 次
- Crucible: Quantifying the Potential of Control Algorithms through LLM AgentsLianchen Jia, Chaoyang Li, Qian Houde, Tianchi Huang 等NeurIPS 2025 · 被引用 1 次
- SimWorld: An Open-ended Simulator for Agents in Physical and Social WorldsXiaokang Ye, Jiawei Ren, Yan Zhuang, Xuhong He 等NeurIPS 2025 · 被引用 8 次
- SimuHome: A Temporal- and Environment-Aware Benchmark for Smart Home LLM AgentsGyuhyeon Seo, Jungwoo Yang, Junseong Pyo, Nalim Kim 等ICLR 2026 · 被引用 14 次
- AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World ContextsKeyu Li, Junhao Shi, Yang Xiao, Mohan Jiang 等ACL 2026 · 被引用 14 次
