DARG: Dynamic Evaluation of Large Language Models via Adaptive Reasoning Graph
Zhehao Zhang, Jiaao Chen, Diyi Yang
摘要
The current paradigm of evaluating Large Language Models (LLMs) through static benchmarks comes with significant limitations, such as vulnerability to data contamination and a lack of adaptability to the evolving capabilities of LLMs. Therefore, evaluation methods that can adapt and generate evaluation data with controlled complexity are urgently needed. In this work, we introduce Dynamic Evaluation of LLMs via Adaptive Reasoning Graph Evolvement (DARG) to dynamically extend current benchmarks with controlled complexity and diversity. Specifically, we first extract the reasoning graphs of data points in current benchmarks and then perturb the reasoning graphs to generate novel testing data. Such newly generated test samples can have different levels of complexity while maintaining linguistic diversity similar to the original benchmarks. We further use a code-augmented LLM to ensure the label correctness of newly generated data. We apply our DARG framework to diverse reasoning tasks in four domains with 15 state-of-the-art LLMs. Experimental results show that almost all LLMs experience a performance decrease with increased complexity and certain LLMs exhibit significant drops. Additionally, we find that LLMs exhibit more biases when being evaluated via the data generated by DARG with higher complexity levels. These observations provide useful insights into how to dynamically and adaptively evaluate LLMs. The code is available at https://github.com/SALT-NLP/DARG.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- NetArena: Dynamic Benchmarks for AI Agents in Network AutomationYajie Zhou, Jiajun Ruan, Eric S. Wang, Sadjad Fouladi 等ICLR 2026 · 被引用 17 次
- Local Success Does Not Compose: Benchmarking Large Language Models for Compositional Formal VerificationXu Xu, Xin Li, Xingwei Qu, Jie Fu 等ICLR 2026 · 被引用 9 次
- CapBencher: Give Your LLM Benchmark a Built-in Alarm for Test-Set OverfittingTakashi Ishida, Thanawat Lodkaew, Ikko YamaneICML 2026 · 被引用 4 次
- Inflated Excellence or True Performance? Rethinking Medical Diagnostic Benchmarks with Dynamic EvaluationXiangxu Zhang, Lei Li, Yanyun Zhou, Xiao Zhou 等ACL 2026 · 被引用 3 次
- Benchmarking Large Language Models Under Data Contamination: A Survey from Static to Dynamic EvaluationSimin Chen, Yiming Chen, Zexin Li, Yifan Jiang 等EMNLP 2025 · 被引用 2 次
它引用的顶会 Paper51
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
相关 Paper
- DyCodeEval: Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data ContaminationSimin Chen, Pranav Pusarla, Baishakhi RayICML 2025
- DyVal: Dynamic Evaluation of Large Language Models for Reasoning TasksKaijie Zhu, Jiaao Chen, Jindong Wang, Neil Zhenqiang Gong 等ICLR 2024 · 被引用 92 次
- Autonomous Evaluation of LLMs for Truth Maintenance and Reasoning TasksRushang Karia, Daniel Bramblett, Daksh Dobhal, Siddharth SrivastavaICLR 2025
- M³GQA: A Multi-Entity Multi-Hop Multi-Setting Graph Question Answering BenchmarkBoci Peng, Yongchao Liu, Xiaohe Bo, Jiaxin Guo 等ACL 2025 · 被引用 1 次
- From Static Benchmarks to Dynamic Protocol: Agent-Centric Text Anomaly Detection for Evaluating LLM ReasoningSeungdong Yoa, Sanghyu Yoon, Suhee Yoon, Dongmin Kim 等ICLR 2026 · 被引用 1 次
