Hierarchical Divide-and-Conquer for Fine-Grained Alignment in LLM-Based Medical Evaluation
Shunfan Zheng, Xiechi Zhang, Gerard de Melo, Xiaoling Wang, Linlin Wang
摘要
In the rapidly evolving landscape of large language models (LLMs) for medical applications, ensuring the reliability and accuracy of these models in clinical settings is paramount. Existing benchmarks often focus on fixed-format tasks like multiple-choice QA, which fail to capture the complexity of real-world clinical diagnostics. Moreover, traditional evaluation metrics and LLM-based evaluators struggle with misalignment, often providing oversimplified assessments that do not adequately reflect human judgment. To address these challenges, we introduce HDCEval 1 , a Hierarchical Divideand-Conquer Evaluation framework tailored for fine-grained alignment in medical evaluation. HDCEval is built on a set of fine-grained medical evaluation guidelines developed in collaboration with professional doctors, encompassing Patient Question Relevance, Medical Knowledge Correctness, and Expression. The framework decomposes complex evaluation tasks into specialized subtasks, each evaluated by expert models trained through Attribute-Driven Token Optimization (ADTO) on a meticulously curated preference dataset. This hierarchical approach ensures that each aspect of the evaluation is handled with expert precision, leading to a significant improvement in alignment with human evaluators.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper9
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning OptimizationYidong Wang, Zhuohao Yu, Wenjin Yao, Zhengran Zeng 等ICLR 2024 · 被引用 368 次
- Asking and Answering Questions to Evaluate the Factual Consistency of SummariesAlex Wang, Kyunghyun Cho, Mike LewisACL 2020 · 被引用 317 次
- MedDialog: Large-scale Medical Dialogue DatasetsGuangtao Zeng, Wenmian Yang, Zeqian Ju, Yue Yang 等EMNLP 2020 · 被引用 163 次
相关 Paper
- HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria DecompositionYuxuan Liu, Tianchi Yang, Shaohan Huang, Zihan Zhang 等ACL 2024 · 被引用 1 次
- AutoMedEval: Harnessing Language Models for Automatic Medical Capability EvaluationXiechi Zhang, Zetian Ouyang, Linlin Wang, Gerard de Melo 等ACL 2025 · 被引用 1 次
- Inflated Excellence or True Performance? Rethinking Medical Diagnostic Benchmarks with Dynamic EvaluationXiangxu Zhang, Lei Li, Yanyun Zhou, Xiao Zhou 等ACL 2026 · 被引用 3 次
- CMedCalc-Bench: A Fine-Grained Benchmark for Chinese Medical Calculations in LLMYunyan Zhang, Zhihong Zhu, Xian WuEMNLP 2025 · 被引用 1 次
- Measuring the Unmeasurable: Unveiling Latent Cognitive Capabilities of LLMCui Danxin, Sihang Jiang, Keyi Wang, Zhiyi Duan 等AAAI 2026
