Evaluating the Knowledge Dependency of Questions
Hyeongdon Moon, Yoonseok Yang, Hangyeol Yu, Seunghyun Lee, Myeongho Jeong, Juneyoung Park, Jamin Shin, Minsam Kim, Seungtaek Choi
Abstract
The automatic generation of Multiple Choice Questions (MCQ) has the potential to reduce the time educators spend on student assessment significantly. However, existing evaluation metrics for MCQ generation, such as BLEU, ROUGE, and METEOR, focus on the n-gram based similarity of the generated MCQ to the gold sample in the dataset and disregard their educational value. They fail to evaluate the MCQ's ability to assess the student's knowledge of the corresponding target fact. To tackle this issue, we propose a novel automatic evaluation metric, coined Knowledge Dependent Answerability (KDA), which measures the MCQ's answerability given knowledge of the target fact. Specifically, we first show how to measure KDA based on student responses from a human survey. Then, we propose two automatic evaluation metrics, KDA disc and KDA cont , that approximate KDA by leveraging pre-trained language models to imitate students' problem-solving behavior. Through our human studies, we show that KDA disc and KDA cont have strong correlations with both (1) KDA and (2) usability in an actual classroom setting, labeled by experts. Furthermore, when combined with ngram based similarity metrics, KDA disc and KDA cont are shown to have a strong predictive power for various expert-labeled MCQ quality measures. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8ece6fcd-5c49-479c-966e-4ea307eb2ef4Cited by top-tier papers4
- Distractor Generation in Multiple-Choice Tasks: A Survey of Methods, Datasets, and EvaluationElaf Alhazmi, Quan Sheng, Wei Emma Zhang, Munazza Zaib et al.EMNLP 2024 · 15 citations
- QG-SMS: Enhancing Test Item Analysis via Student Modeling and SimulationBang Nguyen, Tingting Du, Mengxia Yu, Lawrence Angrave et al.ACL 2025 · 2 citations
- Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the AboveNishant Balepur, Rachel Rudinger, Jordan Lee Boyd-GraberACL 2025
- Generating Plausible Distractors for Multiple-Choice Questions via Student Choice PredictionYooseop Lee, Suin Kim, Yohan JoACL 2025
Builds on4
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu et al.NeurIPS 2020 · 1,957 citations
- Knowledge-Driven Distractor Generation for Cloze-Style Multiple Choice QuestionsSiyu Ren, Kenny Q. ZhuAAAI 2021 · 62 citations
Related papers
- QGEval: Benchmarking Multi-dimensional Evaluation for Question GenerationWeiping Fu, Bifan Wei, Jianxiang Hu, Zhongmin Cai et al.EMNLP 2024 · 8 citations
- RoMe: A Robust Metric for Evaluating Natural Language GenerationMd Rashad Al Hasan Rony, Liubov Kovriguina, Debanjan Chaudhuri, Ricardo Usbeck et al.ACL 2022 · 15 citations
- FEQA: A Question Answering Evaluation Framework for Faithfulness Assessment in Abstractive SummarizationEsin Durmus, He He, Mona T. DiabACL 2020 · 90 citations
- Diversity Enhanced Narrative Question Generation for StorybooksHokeun Yoon, JinYeong BakEMNLP 2023 · 5 citations
- : Evaluating Factual Consistency in Knowledge-Grounded Dialogues via Question Generation and Question AnsweringOr Honovich, Leshem Choshen, Roee Aharoni, Ella Neeman et al.EMNLP 2021 · 101 citations
