QAEval: Mixture of Evaluators for Question-Answering Task Evaluation
Tan Yue, Rui Mao, Xuzhao Shi, Shuo Zhan, Zuhao Yang, Dongyan Zhao
Abstract
Question answering (QA) tasks serve as a key benchmark for evaluating generation systems. Traditional rule-based metrics, such as accuracy and relaxed-accuracy, struggle with open-ended and unstructured responses. LLM-based evaluation methods offer greater flexibility but suffer from sensitivity to instructions, robustness issues, and high computational costs. To overcome these challenges, we introduce QAE-val, a hybrid framework combining rule-based reliability with LLM-based adaptability. QAE-val utilizes two high-quality datasets: QAEx-tract for short-answer extraction and QAScore for scoring model training. By integrating a Mixture of Evaluators model with Dynamic Load Balancing Optimization, QAEval enables accurate, cost-effective QA evaluation. Experimental results show it outperforms models like GPT-4o and Claude-3, achieving 92.3% accuracy with only 0.6B parameters.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 34b4c577-95ec-49c9-b7ae-f731b8d85148Cited by top-tier papers3
- CaptionQA: Is Your Caption as Useful as the Image Itself?Shijia Yang, Yunong Liu, Bohan Zhai, Ximeng Sun et al.CVPR 2026 · 15 citations
- Towards Trustworthy Video Anomaly Understanding: A Class-Guided Chain-of-Evaluation Metric and An Anomaly-focused Meta-BenchmarkJiaxu Leng, Zhoujie Huang, Mingpi Tan, Zhanjie Wu et al.ICML 2026
- MARS: Multimodal Adaptive Reasoning Model for Avoiding OverthinkingTan Yue, Qiong Wu, Dongyan ZhaoAAAI 2026
Builds on8
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language ModelsSeungone Kim, Juyoung Suk, Shayne Longpre, Bill Yuchen Lin et al.EMNLP 2024 · 38 citations
- Tomayto, Tomahto. Beyond Token-level Answer Equivalence for Question Answering EvaluationJannis Bulian, Christian Buck, Wojciech Gajewski, Benjamin Börschinger et al.EMNLP 2022 · 34 citations
- A Critical Evaluation of Evaluations for Long-form Question AnsweringFangyuan Xu, Yixiao Song, Mohit Iyyer, Eunsol ChoiACL 2023 · 25 citations
Related papers
- IQA-EVAL: Automatic Evaluation of Human-Model Interactive Question AnsweringRuosen Li, Ruochen Li, Barry Wang, Xinya DuNeurIPS 2024 · 26 citations
- Auto-PRE: An Automatic and Cost-Efficient Peer-Review Framework for Language Generation EvaluationJunjie Chen, Weihang Su, Zhumin Chu, Haitao Li et al.AAAI 2026
- Learning Answer Generation using Supervision from Automatic Question Answering EvaluatorsMatteo Gabburo, Siddhant Garg, Rik Koncel-Kedziorski, Alessandro MoschittiACL 2023 · 4 citations
- QRelScore: Better Evaluating Generated Questions with Deeper Understanding of Context-aware RelevanceXiaoqiang Wang, Bang Liu, Siliang Tang, Lingfei WuEMNLP 2022 · 6 citations
- MixEval: Deriving Wisdom of the Crowd from LLM Benchmark MixturesJinjie Ni, Fuzhao Xue, Xiang Yue, Yuntian Deng et al.NeurIPS 2024 · 88 citations
