HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition
Yuxuan Liu, Tianchi Yang, Shaohan Huang, Zihan Zhang, Haizhen Huang, Furu Wei, Weiwei Deng, Feng Sun, Qi Zhang
Abstract
Large language models (LLMs) have emerged as a promising alternative to expensive human evaluations. However, the alignment and coverage of LLM-based evaluations are often limited by the scope and potential bias of the evaluation prompts and criteria. To address this challenge, we propose HD-EVAL, a novel framework that iteratively aligns LLM-based evaluators with human preference via Hierarchical Criteria Decomposition. HD-EVAL inherits the essence from the evaluation mindset of human experts and enhances the alignment of LLM-based evaluators by decomposing a given evaluation task into finer-grained criteria, aggregating them according to estimated human preferences, pruning insignificant criteria with attribution, and further decomposing significant criteria. By integrating these steps within an iterative alignment training process, we obtain a hierarchical decomposition of criteria that comprehensively captures aspects of natural language at multiple levels of granularity. Implemented as a white box, the human preferenceguided aggregator is efficient to train and more explainable than relying solely on prompting, and its independence from model parameters makes it applicable to closed-source LLMs. Extensive experiments on three evaluation domains demonstrate the superiority of HD-EVAL in further aligning state-of-the-art evaluators and providing deeper insights into the explanation of evaluation results and the task itself.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d96c6e6c-10af-4c6c-a9b2-3b5d82e73db2Cited by top-tier papers6
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judgeDawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi et al.EMNLP 2025 · 37 citations
- Crowd Comparative Reasoning: Unlocking Comprehensive Evaluations for LLM-as-a-JudgeQiyuan Zhang, Yufei Wang, Yuxin Jiang, Liangyou Li et al.ACL 2025 · 16 citations
- Beyond the Surface: Enhancing LLM-as-a-Judge Alignment with Human via Internal RepresentationsPeng Lai, Jianjie Zheng, Sijie Cheng, Yun Chen et al.NeurIPS 2025 · 16 citations
- FRABench and UFEval: Unified Fine-grained Evaluation with Task and Aspect GeneralizationShibo Hong, Jiahao Ying, Haiyuan Liang, Mengdi Zhang et al.ICLR 2026 · 2 citations
- CheckEval: A reliable LLM-as-a-Judge framework for evaluating text generation using checklistsYukyung Lee, JoongHoon Kim, Jaehee Kim, Hyowon Cho et al.EMNLP 2025 · 2 citations
Builds on14
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- LIMA: Less Is More for AlignmentChunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer et al.NeurIPS 2023 · 1,486 citations
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 1,143 citations
- ChatEval: Towards Better LLM-based Evaluators through Multi-Agent DebateChi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu et al.ICLR 2024 · 871 citations
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang et al.EMNLP 2023 · 549 citations
Related papers
- Hierarchical Divide-and-Conquer for Fine-Grained Alignment in LLM-Based Medical EvaluationShunfan Zheng, Xiechi Zhang, Gerard de Melo, Xiaoling Wang et al.AAAI 2025 · 4 citations
- EvalLM: Interactive Evaluation of Large Language Model Prompts on User-Defined CriteriaTae Soo Kim, Yoonjoo Lee, Jamin Shin, Young-Ho Kim et al.CHI 2024 · 81 citations
- Co-Eval: Augmenting LLM-based Evaluation with Machine MetricsLing-I Wu, Weijie Wu, Minyu Chen, Jianxin Xue et al.EMNLP 2025
- HypoEval: Hypothesis-Guided Evaluation for Natural Language GenerationMingxuan Li, Hanchen Li, Chenhao TanACL 2026 · 1 citation
- Multi-Domain Explainability of PreferencesNitay Calderon, Liat Ein-Dor, Roi ReichartEMNLP 2025
