Evaluating the Knowledge Dependency of Questions
Hyeongdon Moon, Yoonseok Yang, Hangyeol Yu, Seunghyun Lee, Myeongho Jeong, Juneyoung Park, Jamin Shin, Minsam Kim, Seungtaek Choi
摘要
The automatic generation of Multiple Choice Questions (MCQ) has the potential to reduce the time educators spend on student assessment significantly. However, existing evaluation metrics for MCQ generation, such as BLEU, ROUGE, and METEOR, focus on the n-gram based similarity of the generated MCQ to the gold sample in the dataset and disregard their educational value. They fail to evaluate the MCQ's ability to assess the student's knowledge of the corresponding target fact. To tackle this issue, we propose a novel automatic evaluation metric, coined Knowledge Dependent Answerability (KDA), which measures the MCQ's answerability given knowledge of the target fact. Specifically, we first show how to measure KDA based on student responses from a human survey. Then, we propose two automatic evaluation metrics, KDA disc and KDA cont , that approximate KDA by leveraging pre-trained language models to imitate students' problem-solving behavior. Through our human studies, we show that KDA disc and KDA cont have strong correlations with both (1) KDA and (2) usability in an actual classroom setting, labeled by experts. Furthermore, when combined with ngram based similarity metrics, KDA disc and KDA cont are shown to have a strong predictive power for various expert-labeled MCQ quality measures. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Distractor Generation in Multiple-Choice Tasks: A Survey of Methods, Datasets, and EvaluationElaf Alhazmi, Quan Sheng, Wei Emma Zhang, Munazza Zaib 等EMNLP 2024 · 被引用 15 次
- QG-SMS: Enhancing Test Item Analysis via Student Modeling and SimulationBang Nguyen, Tingting Du, Mengxia Yu, Lawrence Angrave 等ACL 2025 · 被引用 2 次
- Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the AboveNishant Balepur, Rachel Rudinger, Jordan Lee Boyd-GraberACL 2025
- Generating Plausible Distractors for Multiple-Choice Questions via Student Choice PredictionYooseop Lee, Suin Kim, Yohan JoACL 2025
它引用的顶会 Paper4
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu 等NeurIPS 2020 · 被引用 1,957 次
- Knowledge-Driven Distractor Generation for Cloze-Style Multiple Choice QuestionsSiyu Ren, Kenny Q. ZhuAAAI 2021 · 被引用 62 次
相关 Paper
- QGEval: Benchmarking Multi-dimensional Evaluation for Question GenerationWeiping Fu, Bifan Wei, Jianxiang Hu, Zhongmin Cai 等EMNLP 2024 · 被引用 8 次
- RoMe: A Robust Metric for Evaluating Natural Language GenerationMd Rashad Al Hasan Rony, Liubov Kovriguina, Debanjan Chaudhuri, Ricardo Usbeck 等ACL 2022 · 被引用 15 次
- FEQA: A Question Answering Evaluation Framework for Faithfulness Assessment in Abstractive SummarizationEsin Durmus, He He, Mona T. DiabACL 2020 · 被引用 90 次
- Diversity Enhanced Narrative Question Generation for StorybooksHokeun Yoon, JinYeong BakEMNLP 2023 · 被引用 5 次
- : Evaluating Factual Consistency in Knowledge-Grounded Dialogues via Question Generation and Question AnsweringOr Honovich, Leshem Choshen, Roee Aharoni, Ella Neeman 等EMNLP 2021 · 被引用 101 次
