Beyond Benchmarks: Toward Causally Faithful Evaluation of Large Language Models
Zhengshuyuan Tian, Wanling Gao, Chuanxin Lan, Chenxi Wang, Lei Wang, Guoxin Kang, Zhengxin Yang, Yunyou Huang, Xuehai Hong, Jianfeng Zhan
Abstract
Current LLM evaluations often conflate benchmark performance with intrinsic model capability. This is misleading, as observed outcomes arise from the entire evaluation system, including datasets, prompting methods, decoding parameters, and the software–hardware stack, rather than the model alone. When this system is under-specified, attribution becomes unreliable; in practice, evaluation choices alone can induce accuracy swings of up to 70%. This challenge is compounded by the open-ended nature of LLM evaluation, where questions span languages, domains, and usage styles, forming variable and implicitly shifting datasets. Consequently, strong performance on static benchmarks may reflect surface alignment or dataset-induced effects rather than robust capability. Prior studies often focus on individual components or manually-curated small-scale dataset variants, overlooking interactions and dataset-related confounding. To address these limitations, we propose LLM evaluatology, a principled framework that grounds LLM evaluation in a causally motivated system design. It combines structured causal modeling as an intervention-oriented lens with factorial decomposition under design of experiments, quantifying main and interaction effects while using instance-level interventions to probe dataset-induced effects. By jointly modeling evaluation components and structured question variations, LLM evaluatology enables more interpretable, reproducible, and carefully attributed assessment of model capability. Our framework is publicly available at GitHub .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ebc1fa8c-c1ce-411d-900a-ee8804466c5cBuilds on17
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu et al.ICLR 2024 · 1,472 citations
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formattingMelanie Sclar, Yejin Choi, Yulia Tsvetkov, Alane SuhrICLR 2024 · 682 citations
Related papers
- Can Prompt Probe Pretrained Language Models? Understanding the Invisible Risks from a Causal ViewBoxi Cao, Hongyu Lin, Xianpei Han, Fangchao Liu et al.ACL 2022
- Unbiased Evaluation of Large Language Models from a Causal PerspectiveMeilin Chen, Jian Tian, Liang Ma, Di Xie et al.ICML 2025
- NoisyCausal: A Benchmark for Evaluating Causal Reasoning Under Structured NoiseZhi Xu, Yun FuACL 2026
- Measuring Intent Comprehension in LLMsNadav Kunievsky, James EvansICML 2026 · 1 citation
- Bridging Internal Consistency and External Alignment: A Causal and Dynamic Interpretability Framework for LLM GenerationShuyao Xiao, Shengling Wang, Ke ChaoACL 2026
