LIQUID: A Framework for List Question Answering Dataset Generation
Seongyun Lee, Hyunjae Kim, Jaewoo Kang
Abstract
Question answering (QA) models often rely on large-scale training datasets, which necessitates the development of a data generation framework to reduce the cost of manual annotations. Although several recent studies have aimed to generate synthetic questions with single-span answers, no study has been conducted on the creation of list questions with multiple, non-contiguous spans as answers. To address this gap, we propose LIQUID, an automated framework for generating list QA datasets from unlabeled corpora. We first convert a passage from Wikipedia or PubMed into a summary and extract named entities from the summarized text as candidate answers. This allows us to select answers that are semantically correlated in context and is, therefore, suitable for constructing list questions. We then create questions using an off-the-shelf question generator with the extracted entities and original passage. Finally, iterative filtering and answer expansion are performed to ensure the accuracy and completeness of the answers. Using our synthetic data, we significantly improve the performance of the previous best list QA models by exact-match F1 scores of 5.0 on MultiSpanQA, 1.9 on Quoref, and 2.8 averaged across three BioASQ benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e91dda76-d703-4dfb-a21a-1b91c4d54561Cited by top-tier papers4
- A Cooperative Multi-Agent Framework for Zero-Shot Named Entity RecognitionZihan Wang, Ziqi Zhao, Yougang Lyu, Zhumin Chen et al.WWW 2025 · 16 citations
- Enhancing Pre-Trained Generative Language Models with Question Attended Span Extraction on Machine Reading ComprehensionLin Ai, Zheng Hui, Zizhou Liu, Julia HirschbergEMNLP 2024 · 4 citations
- Knowledge-to-Verification: Exploring RLVR for LLMs in Knowledge-Intensive DomainsZhonghang Yuan, Zhefan Wang, Fang Hu, Zihong Chen et al.ACL 2026
- Auto-GDA: Automatic Domain Adaptation for Efficient Grounding Verification in Retrieval-Augmented GenerationTobias Leemann, Periklis Petridis, Giuseppe Vietri, Dionysis Manousakas et al.ICLR 2025
Builds on5
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Coreferential Reasoning Learning for Language RepresentationDeming Ye, Yankai Lin, Jiaju Du, Zhenghao Liu et al.EMNLP 2020 · 164 citations
- End-to-End Synthetic Data Generation for Domain Adaptation of Question Answering SystemsSiamak Shakeri, Cícero Nogueira dos Santos, Henghui Zhu, Patrick Ng et al.EMNLP 2020 · 60 citations
- Improving Unsupervised Question Answering via Summarization-Informed Question GenerationChenyang Lyu, Lifeng Shang, Yvette Graham, Jennifer Foster et al.EMNLP 2021 · 33 citations
- Training Question Answering Models From Synthetic DataRaul Puri, Ryan Spring, Mohammad Shoeybi, Mostofa Patwary et al.EMNLP 2020 · 15 citations
Related papers
- Achieving Human Parity in Content-Grounded Datasets GenerationAsaf Yehudai, Boaz Carmeli, Yosi Mass, Ofir Arviv et al.ICLR 2024 · 9 citations
- Generating Information-Seeking Conversations from Unlabeled DocumentsGangwoo Kim, Sungdong Kim, Kang Min Yoo, Jaewoo KangEMNLP 2022 · 4 citations
- Harvesting and Refining Question-Answer Pairs for Unsupervised QAZhongli Li, Wenhui Wang, Li Dong, Furu Wei et al.ACL 2020 · 29 citations
- ASQA: Factoid Questions Meet Long-Form AnswersIvan Stelmakh, Yi Luan, Bhuwan Dhingra, Ming-Wei ChangEMNLP 2022 · 51 citations
- Knowledge Transfer from Answer Ranking to Answer GenerationMatteo Gabburo, Rik Koncel-Kedziorski, Siddhant Garg, Luca Soldaini et al.EMNLP 2022 · 4 citations
