QRelScore: Better Evaluating Generated Questions with Deeper Understanding of Context-aware Relevance
Xiaoqiang Wang, Bang Liu, Siliang Tang, Lingfei Wu
摘要
Existing metrics for assessing question generation not only require costly human reference but also fail to take into account the input context of generation, rendering the lack of deep understanding of the relevance between the generated questions and input contexts. As a result, they may wrongly penalize a legitimate and reasonable candidate question when it (i) involves complicated reasoning with the context or (ii) can be grounded by multiple evidences in the context. In this paper, we propose QRelScore, a context-aware Relevance evaluation metric for Question Generation. Based on off-the-shelf language models such as BERT and GPT2, QRelScore employs both word-level hierarchical matching and sentence-level prompt-based generation to cope with the complicated reasoning and diverse generation from multiple evidences, respectively. Compared with existing metrics, our experiments demonstrate that QRelScore is able to achieve a higher correlation with human judgments while being much more robust to adversarial samples.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- QGEval: Benchmarking Multi-dimensional Evaluation for Question GenerationWeiping Fu, Bifan Wei, Jianxiang Hu, Zhongmin Cai 等EMNLP 2024 · 被引用 8 次
- Dialogizer: Context-aware Conversational-QA Dataset Generation from Textual SourcesYerin Hwang, Yongil Kim, Hyunkyung Bae, Hwanhee Lee 等EMNLP 2023 · 被引用 3 次
- MP2D: An Automated Topic Shift Dialogue Generation Framework Leveraging Knowledge GraphsYerin Hwang, Yongil Kim, Yunah Jang, Jeesoo Bang 等EMNLP 2024 · 被引用 2 次
它引用的顶会 Paper20
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 被引用 1,143 次
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras 等EMNLP 2021 · 被引用 937 次
- Towards Understanding and Mitigating Social Biases in Language ModelsPaul Pu Liang, Chiyu Wu, Louis-Philippe Morency, Ruslan SalakhutdinovICML 2021 · 被引用 495 次
相关 Paper
- Language Model Augmented Relevance ScoreRuibo Liu, Jason Wei, Soroush VosoughiACL 2021
- Improving Image Captioning Evaluation by Considering Inter References VarianceYanzhi Yi, Hangyu Deng, Jinglu HuACL 2020 · 被引用 44 次
- Contextualized Sequence Likelihood: Enhanced Confidence Scores for Natural Language GenerationZhen Lin, Shubhendu Trivedi, Jimeng SunEMNLP 2024 · 被引用 1 次
- TabReX: Tabular Referenceless eXplainable EvaluationTejas Anvekar, Junha Park, Aparna Garimella, Vivek GuptaACL 2026
- Global Explainability of BERT-Based Evaluation Metrics by Disentangling along Linguistic FactorsMarvin Kaster, Wei Zhao, Steffen EgerEMNLP 2021 · 被引用 13 次
