Training Question Answering Models From Synthetic Data
Raul Puri, Ryan Spring, Mohammad Shoeybi, Mostofa Patwary, Bryan Catanzaro
摘要
Question and answer generation is a data augmentation method that aims to improve question answering (QA) models given the limited amount of human labeled data. However, a considerable gap remains between synthetic and human-generated question-answer pairs. This work aims to narrow this gap by taking advantage of large language models and explores several factors such as model size, quality of pretrained models, scale of data synthesized, and algorithmic choices. On the SQUAD1.1 question answering task, we achieve higher accuracy using solely synthetic questions and answers than when using the SQUAD1.1 training set questions alone. Removing access to real Wikipedia data, we synthesize questions and answers from a synthetic text corpus generated by an 8.3 billion parameter GPT-2 model and achieve 88.4 Exact Match (EM) and 93.9 F1 score on the SQUAD1.1 dev set. We further apply our methodology to SQUAD2.0 and show a 2.8 absolute gain on EM score compared to prior work using synthetic data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper30
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- ZeroGen: Efficient Zero-shot Learning via Dataset GenerationJiacheng Ye, Jiahui Gao, Qintong Li, Hang Xu 等EMNLP 2022 · 被引用 96 次
- Improving Question Answering Model Robustness with Synthetic Adversarial Data GenerationMax Bartolo, Tristan Thrush, Robin Jia, Sebastian Riedel 等EMNLP 2021 · 被引用 68 次
- End-to-End Synthetic Data Generation for Domain Adaptation of Question Answering SystemsSiamak Shakeri, Cícero Nogueira dos Santos, Henghui Zhu, Patrick Ng 等EMNLP 2020 · 被引用 60 次
- Turning Tables: Generating Examples from Semi-structured Tables for Endowing Language Models with Reasoning SkillsOri Yoran, Alon Talmor, Jonathan BerantACL 2022 · 被引用 57 次
它引用的顶会 Paper3
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Large Scale Multi-Actor Generative Dialog ModelingAlex Boyd, Raul Puri, Mohammad Shoeybi, Mostofa Patwary 等ACL 2020 · 被引用 1 次
- Self-Training With Noisy Student Improves ImageNet ClassificationQizhe Xie, Minh-Thang Luong, Eduard H. Hovy, Quoc V. LeCVPR 2020
相关 Paper
- Harvesting and Refining Question-Answer Pairs for Unsupervised QAZhongli Li, Wenhui Wang, Li Dong, Furu Wei 等ACL 2020 · 被引用 29 次
- Improving Unsupervised Question Answering via Summarization-Informed Question GenerationChenyang Lyu, Lifeng Shang, Yvette Graham, Jennifer Foster 等EMNLP 2021 · 被引用 33 次
- Leveraging QA Datasets to Improve Generative Data AugmentationDheeraj Mekala, Tu Vu, Timo Schick, Jingbo ShangEMNLP 2022 · 被引用 8 次
- Data-Centric Lessons To Improve Speech-Language PretrainingVishaal Udandarao, Zhiyun Lu, Xuankai Chang, Yongqiang Wang 等ICLR 2026 · 被引用 3 次
- Improving Question Generation with Sentence-Level Semantic Matching and Answer Position InferringXiyao Ma, Qile Zhu, Yanlin Zhou, Xiaolin LiAAAI 2020 · 被引用 68 次
