MetaMetrics: Calibrating Metrics for Generation Tasks Using Human Preferences
Genta Indra Winata, David Anugraha, Lucky Susanto, Garry Kuwanto, Derry Tanti Wijaya
摘要
Understanding the quality of a performance evaluation metric is crucial for ensuring that model outputs align with human preferences. However, it remains unclear how well each metric captures the diverse aspects of these preferences, as metrics often excel in one particular area but not across all dimensions. To address this, it is essential to systematically calibrate metrics to specific aspects of human preference, catering to the unique characteristics of each aspect. We introduce METAMETRICS, a calibrated meta-metric designed to evaluate generation tasks across different modalities in a supervised manner. METAMETRICS optimizes the combination of existing metrics to enhance their alignment with human preferences. Our metric demonstrates flexibility and effectiveness in both language and vision downstream tasks, showing significant benefits across various multilingual and multi-domain scenarios. METAMETRICS aligns closely with human preferences and is highly extendable and easily integrable into any application. This makes METAMETRICS a powerful tool for improving the evaluation of generation tasks, ensuring that metrics are more representative of human judgment across diverse contexts. * The work was done outside Capital One. † Equal contribution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Reward Reasoning ModelsJiaxin Guo, Zewen Chi, Li Dong, Qingxiu Dong 等NeurIPS 2025 · 被引用 14 次
- AutoMetrics: Approximate Human Judgments with Automatically Generated EvaluatorsMichael J Ryan, Yanzhe Zhang, Amol Salunkhe, Yi Chu 等ICLR 2026 · 被引用 3 次
- Bridging Internal Consistency and External Alignment: A Causal and Dynamic Interpretability Framework for LLM GenerationShuyao Xiao, Shengling Wang, Ke ChaoACL 2026
它引用的顶会 Paper10
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 被引用 1,143 次
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras 等EMNLP 2021 · 被引用 937 次
相关 Paper
- NLG Evaluation Metrics Beyond Correlation Analysis: An Empirical Metric Preference ChecklistIftitahu Ni'mah, Meng Fang, Vlado Menkovski, Mykola PechenizkiyACL 2023 · 被引用 9 次
- Learning Multi-Dimensional Human Preference for Text-to-Image GenerationSixian Zhang, Bohan Wang, Junqiang Wu, Yan Li 等CVPR 2024 · 被引用 15 次
- COMET: A Neural Framework for MT EvaluationRicardo Rei, Craig Stewart, Ana C. Farinha, Alon LavieEMNLP 2020 · 被引用 6 次
- Reuse Your Rewards: Reward Model Transfer for Zero-Shot Cross-Lingual AlignmentZhaofeng Wu, Ananth Balashankar, Yoon Kim, Jacob Eisenstein 等EMNLP 2024 · 被引用 29 次
- Co-Eval: Augmenting LLM-based Evaluation with Machine MetricsLing-I Wu, Weijie Wu, Minyu Chen, Jianxin Xue 等EMNLP 2025
