Training Question Answering Models From Synthetic Data
Raul Puri, Ryan Spring, Mohammad Shoeybi, Mostofa Patwary, Bryan Catanzaro
Abstract
Question and answer generation is a data augmentation method that aims to improve question answering (QA) models given the limited amount of human labeled data. However, a considerable gap remains between synthetic and human-generated question-answer pairs. This work aims to narrow this gap by taking advantage of large language models and explores several factors such as model size, quality of pretrained models, scale of data synthesized, and algorithmic choices. On the SQUAD1.1 question answering task, we achieve higher accuracy using solely synthetic questions and answers than when using the SQUAD1.1 training set questions alone. Removing access to real Wikipedia data, we synthesize questions and answers from a synthetic text corpus generated by an 8.3 billion parameter GPT-2 model and achieve 88.4 Exact Match (EM) and 93.9 F1 score on the SQUAD1.1 dev set. We further apply our methodology to SQUAD2.0 and show a 2.8 absolute gain on EM score compared to prior work using synthetic data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f4367b1c-a5c8-4128-8d0b-08fa8402a60cCited by top-tier papers30
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- ZeroGen: Efficient Zero-shot Learning via Dataset GenerationJiacheng Ye, Jiahui Gao, Qintong Li, Hang Xu et al.EMNLP 2022 · 96 citations
- Improving Question Answering Model Robustness with Synthetic Adversarial Data GenerationMax Bartolo, Tristan Thrush, Robin Jia, Sebastian Riedel et al.EMNLP 2021 · 68 citations
- End-to-End Synthetic Data Generation for Domain Adaptation of Question Answering SystemsSiamak Shakeri, Cícero Nogueira dos Santos, Henghui Zhu, Patrick Ng et al.EMNLP 2020 · 60 citations
- Turning Tables: Generating Examples from Semi-structured Tables for Endowing Language Models with Reasoning SkillsOri Yoran, Alon Talmor, Jonathan BerantACL 2022 · 57 citations
Builds on3
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Large Scale Multi-Actor Generative Dialog ModelingAlex Boyd, Raul Puri, Mohammad Shoeybi, Mostofa Patwary et al.ACL 2020 · 1 citation
- Self-Training With Noisy Student Improves ImageNet ClassificationQizhe Xie, Minh-Thang Luong, Eduard H. Hovy, Quoc V. LeCVPR 2020
Related papers
- Harvesting and Refining Question-Answer Pairs for Unsupervised QAZhongli Li, Wenhui Wang, Li Dong, Furu Wei et al.ACL 2020 · 29 citations
- Improving Unsupervised Question Answering via Summarization-Informed Question GenerationChenyang Lyu, Lifeng Shang, Yvette Graham, Jennifer Foster et al.EMNLP 2021 · 33 citations
- Leveraging QA Datasets to Improve Generative Data AugmentationDheeraj Mekala, Tu Vu, Timo Schick, Jingbo ShangEMNLP 2022 · 8 citations
- Data-Centric Lessons To Improve Speech-Language PretrainingVishaal Udandarao, Zhiyun Lu, Xuankai Chang, Yongqiang Wang et al.ICLR 2026 · 3 citations
- Improving Question Generation with Sentence-Level Semantic Matching and Answer Position InferringXiyao Ma, Qile Zhu, Yanlin Zhou, Xiaolin LiAAAI 2020 · 68 citations
