On the Generation of Medical Question-Answer Pairs
Sheng Shen, Yaliang Li, Nan Du, Xian Wu, Yusheng Xie, Shen Ge, Tao Yang, Kai Wang, Xingzheng Liang, Wei Fan
Abstract
Question answering (QA) has achieved promising progress recently. However, answering a question in real-world scenarios like the medical domain is still challenging, due to the requirement of external knowledge and the insufficient quantity of high-quality training data. In the light of these challenges, we study the task of generating medical QA pairs in this paper. With the insight that each medical question can be considered as a sample from the latent distribution of questions given answers, we propose an automated medical QA pair generation framework, consisting of an unsupervised key phrase detector that explores unstructured material for validity, and a generator that involves a multi-pass decoder to integrate structural knowledge for diversity. A series of experiments have been conducted on a real-world dataset collected from the National Medical Licensing Examination of China. Both automatic evaluation and human annotation demonstrate the effectiveness of the proposed method. Further investigation shows that, by incorporating the generated QA pairs for training, significant improvement in terms of accuracy can be achieved for the examination QA system. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itRelated papers
- Asking Questions Like Educational Experts: Automatically Generating Question-Answer Pairs on Real-World Examination DataFanyi Qu, Xin Jia, Yunfang WuEMNLP 2021 · 26 citations
- Generating Diverse and Consistent QA pairs from Contexts with Information-Maximizing Hierarchical Conditional VAEsDong Bok Lee, Seanie Lee, Woo Tae Jeong, Donghwan Kim et al.ACL 2020 · 22 citations
- Towards Medical Machine Reading Comprehension with Structural Knowledge and Plain TextDongfang Li, Baotian Hu, Qingcai Chen, Weihua Peng et al.EMNLP 2020 · 39 citations
- MLEC-QA: A Chinese Multi-Choice Biomedical Question Answering DatasetJing Li, Shangping Zhong, Kaizhi ChenEMNLP 2021 · 24 citations
- To Generate or to Retrieve? On the Effectiveness of Artificial Contexts for Medical Open-Domain Question AnsweringGiacomo Frisoni, Alessio Cocchieri, Alex Presepi, Gianluca Moro et al.ACL 2024 · 8 citations
