Synthetic Question Value Estimation for Domain Adaptation of Question Answering
Xiang Yue, Ziyu Yao, Huan Sun
Abstract
Synthesizing QA pairs with a question generator (QG) on the target domain has become a popular approach for domain adaptation of question answering (QA) models. Since synthetic questions are often noisy in practice, existing work adapts scores from a pretrained QA (or QG) model as criteria to select highquality questions. However, these scores do not directly serve the ultimate goal of improving QA performance on the target domain. In this paper, we introduce a novel idea of training a question value estimator (QVE) that directly estimates the usefulness of synthetic questions for improving the target-domain QA performance. By conducting comprehensive experiments, we show that the synthetic questions selected by QVE can help achieve better target-domain QA performance, in comparison with existing techniques. We additionally show that by using such questions and only around 15% of the human annotations on the target domain, we can achieve comparable performance to the fully-supervised baselines. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- ReasTAP: Injecting Table Reasoning Skills During Pre-training via Synthetic Reasoning ExamplesYilun Zhao, Linyong Nan, Zhenting Qi, Rui Zhang et al.EMNLP 2022 · 15 citations
- Modeling What-to-ask and How-to-ask for Answer-unaware Conversational Question GenerationXuan Long Do, Bowei Zou, Shafiq R. Joty, Anh Tran Tai et al.ACL 2023 · 1 citation
- QA Domain Adaptation using Hidden Space Augmentation and Self-Supervised Contrastive AdaptationZhenrui Yue, Huimin Zeng, Bernhard Kratzwald, Stefan Feuerriegel et al.EMNLP 2022 · 1 citation
Builds on11
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Data Valuation using Reinforcement LearningJinsung Yoon, Sercan Ömer Arik, Tomas PfisterICML 2020 · 236 citations
- Asking Questions the Human Way: Scalable Question-Answer Generation from Text CorpusBang Liu, Haojie Wei, Di Niu, Haolan Chen et al.WWW 2020 · 100 citations
- Capturing Greater Context for Question GenerationLuu Anh Tuan, Darsh J. Shah, Regina BarzilayAAAI 2020 · 77 citations
- End-to-End Synthetic Data Generation for Domain Adaptation of Question Answering SystemsSiamak Shakeri, Cícero Nogueira dos Santos, Henghui Zhu, Patrick Ng et al.EMNLP 2020 · 60 citations
Related papers
- Contrastive Domain Adaptation for Question Answering using Limited Text CorporaZhenrui Yue, Bernhard Kratzwald, Stefan FeuerriegelEMNLP 2021 · 22 citations
- Robust Domain Adaptation for Machine Reading ComprehensionLiang Jiang, Zhenyu Huang, Jia Liu, Zujie Wen et al.AAAI 2023 · 1 citation
- Bridging the Synthetic-to-Authentic Gap: Distortion-Guided Unsupervised Domain Adaptation for Blind Image Quality AssessmentAobo Li, Jinjian Wu, Yongxu Liu, Leida LiCVPR 2024
- Improving Unsupervised Question Answering via Summarization-Informed Question GenerationChenyang Lyu, Lifeng Shang, Yvette Graham, Jennifer Foster et al.EMNLP 2021 · 33 citations
- Generating Information-Seeking Conversations from Unlabeled DocumentsGangwoo Kim, Sungdong Kim, Kang Min Yoo, Jaewoo KangEMNLP 2022 · 4 citations
