Achieving Human Parity in Content-Grounded Datasets Generation
Asaf Yehudai, Boaz Carmeli, Yosi Mass, Ofir Arviv, Nathaniel Mills, Eyal Shnarch, Leshem Choshen
Abstract
The lack of high-quality data for content-grounded generation tasks has been identified as a major obstacle to advancing these tasks. To address this gap, we propose Genie, a novel method for automatically generating high-quality contentgrounded data. It consists of three stages: (a) Content Preparation, (b) Generation: creating task-specific examples from the content (e.g., question-answer pairs or summaries). (c) Filtering mechanism aiming to ensure the quality and faithfulness of the generated data. We showcase this methodology by generating three large-scale synthetic data, making wishes, for Long-Form Question-Answering (LFQA), summarization, and information extraction. In a human evaluation, our generated data was found to be natural and of high quality. Furthermore, we compare models trained on our data with models trained on human-written data -ELI5 and ASQA for LFQA and CNN-DailyMail for Summarization. We show that our models are on par with or outperforming models trained on human-generated data and consistently outperforming them in faithfulness. Finally, we applied our method to create LFQA data within the medical domain and compared a model trained on it with models trained on other domains.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f7bdf8dc-3496-42ce-ae02-d7be25def8dfCited by top-tier papers8
- Self-Adapting Language ModelsAdam Zweiger, Jyothish Pari, Han Guo, Yoon Kim et al.NeurIPS 2025 · 78 citations
- ROC-n-reroll: How verifier imperfection affects test-time scalingFlorian E. Dorner, Yatong Chen, André F Cruz, Fanny YangICLR 2026 · 13 citations
- LIFT: A Novel Framework for Enhancing Long-Context Understanding of LLMs via Long Input Fine-TuningYansheng Mao, Yufei Xu, Jiaqi Li, Fanxu Meng et al.ICML 2026 · 7 citations
- Mind the Gap: Examining the Self-Improvement Capabilities of Large Language ModelsYuda Song, Hanlin Zhang, Carson Eisenach, Sham M. Kakade et al.ICLR 2025 · 3 citations
- Knowledge is Not Enough: Injecting RL Skills for Continual AdaptationPingzhi Tang, Yiding Wang, Muhan ZhangACL 2026 · 2 citations
Builds on13
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation DatasetLianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li et al.ICLR 2024 · 419 citations
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 394 citations
Related papers
- LIQUID: A Framework for List Question Answering Dataset GenerationSeongyun Lee, Hyunjae Kim, Jaewoo KangAAAI 2023 · 31 citations
- A Synthetic Data Generation Framework for Grounded DialoguesJianzhu Bao, Rui Wang, Yasheng Wang, Aixin Sun et al.ACL 2023 · 11 citations
- ASQA: Factoid Questions Meet Long-Form AnswersIvan Stelmakh, Yi Luan, Bhuwan Dhingra, Ming-Wei ChangEMNLP 2022 · 51 citations
- Socratic Pretraining: Question-Driven Pretraining for Controllable SummarizationArtidoro Pagnoni, Alexander R. Fabbri, Wojciech Kryscinski, Chien-Sheng WuACL 2023 · 4 citations
- ProxyQA: An Alternative Framework for Evaluating Long-Form Text Generation with Large Language ModelsHaochen Tan, Zhijiang Guo, Zhan Shi, Lu Xu et al.ACL 2024
