AutoTaskEval: Towards Domain-Specific and Fine-Grained Evaluation for LLMs
Qingqing Lyu, Linjuan Wu, Yongliang Shen, Hengwei Liu, Hao Li, Shengpei Jiang, Yin Zhang, Weiming Lu
Abstract
Despite the rapid progress of LLMs, their evaluation remains hindered by static, manually curated benchmarks with limited task coverage and poor adaptability to emerging domains. Existing automated approaches typically operate within fixed task schemas and often fail to autonomously discover new evaluation dimensions, limiting both scalability and effectiveness. To address these gaps, we propose AUTOTASKEVAL, an automated framework that constructs domain-specific benchmarks directly from unstructured corpora. Using a refined Bloom's Taxonomy, the framework systematically discovers tasks, enriches contextual grounding via iterative Socratic prompting, and generates diverse, progressively challenging evaluation instances. Applied to the complex and knowledge-intensive legal domain, AUTO-TASKEVAL uncovers a broader and more finegrained task space than expert-curated benchmarks while producing high-quality instances that preserve established model-level evaluation trends. We further validate its robustness in a low-structure e-commerce review domain. Together, these results show that AU-TOTASKEVAL enables scalable, adaptive, and high-fidelity LLM assessment across domains and model families, advancing autonomous and capability-sensitive evaluation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 12a841c3-7e00-44b9-b468-b12daed75c5cBuilds on4
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- LawBench: Benchmarking Legal Knowledge of Large Language ModelsZhiwei Fei, Xiaoyu Shen, Dawei Zhu, Fengzhe Zhou et al.EMNLP 2024 · 59 citations
- The Art of SOCRATIC QUESTIONING: Recursive Thinking with Large Language ModelsJingyuan Qi, Zhiyang Xu, Ying Shen, Minqian Liu et al.EMNLP 2023 · 20 citations
- AutoBencher: Towards Declarative Benchmark ConstructionXiang Lisa Li, Farzaan Kaiyom, Evan Zheran Liu, Yifan Mai et al.ICLR 2025
Related papers
- Bloom-Eval: A Hierarchical Evaluation Benchmark for Automatic Survey Generation Based on Bloom's TaxonomyFei Zhang, Zhe Zhao, Haibin Wen, Tianshuo Wei et al.ACL 2026
- BiasScope: Towards Automated Detection of Bias in LLM-as-a-Judge EvaluationPeng Lai, Zhihao Ou, Yong Wang, Longyue Wang et al.ICLR 2026 · 15 citations
- SciCustom: A Framework for Custom Evaluation of Scientific Capabilities in Large Language ModelsYiyang Gu, Junwei Yang, Junyu Luo, Ye Yuan et al.ACL 2026
- LegalAgentBench: Evaluating LLM Agents in Legal DomainHaitao Li, Junjie Chen, Jingli Yang, Qingyao Ai et al.ACL 2025
- PLAWBENCH: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal PracticeYuzhen Shi, Huanghai Liu, Yiran Hu, Gaojie Song et al.ACL 2026 · 7 citations
