Generating Biographies on Wikipedia: The Impact of Gender Bias on the Retrieval-Based Generation of Women Biographies
Angela Fan, Claire Gardent
摘要
Generating factual, long-form text such as Wikipedia articles raises three key challenges: how to gather relevant evidence, how to structure information into well-formed text, and how to ensure that the generated text is factually correct. We address these by developing a model for English text that uses a retrieval mechanism to identify relevant supporting information on the web and a cache-based pre-trained encoderdecoder to generate long-form biographies section by section, including citation information. To assess the impact of available web evidence on the output text, we compare the performance of our approach when generating biographies about women (for which less information is available on the web) vs. biographies generally. To this end, we curate a dataset of 1,500 biographies about women. We analyze our generated text to understand how differences in available web evidence data affect generation. We evaluate the factuality, fluency, and quality of the generated texts using automatic metrics and human evaluation. We hope that these techniques can be used as a starting point for human writers, to aid in reducing the complexity inherent in the creation of long-form, factual text.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- The "Colonial Impulse" of Natural Language Processing: An Audit of Bengali Sentiment Analysis Tools and Their Identity-based BiasesDipto Das, Shion Guha, Jed R. Brubaker, Bryan C. SemaanCHI 2024 · 被引用 15 次
- Descartes: Generating Short Descriptions of Wikipedia ArticlesMarija Sakota, Maxime Peyrard, Robert WestWWW 2023 · 被引用 6 次
- Leveraging Outline-Optimized Generative Interactions and Critique for Self-Refining Outlines with Reinforcement LearningHengwei Liu, Haoyuan Ma, Qingqing Lyu, Daoxin Zhang 等ACL 2026
- WikiREVIEW: A Multi-Perspective Review Framework for Automatic Wiki-Style Article GenerationGuo-Biao Zhang, Zhijing Wu, Tian Lan, Ding-Yuan Liu 等AAAI 2026
- Machines in the Margins: A Systematic Review of Automated Content Generation for WikipediaNeal Reeves, Elena SimperlCSCW 2025
它引用的顶会 Paper18
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat 等ICML 2020 · 被引用 2,937 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 被引用 695 次
- Scalable Zero-shot Entity Linking with Dense Entity RetrievalLedell Wu, Fabio Petroni, Martin Josifoski, Sebastian Riedel 等EMNLP 2020 · 被引用 336 次
相关 Paper
- Open Domain Event Text GenerationZihao Fu, Lidong Bing, Wai LamAAAI 2020 · 被引用 9 次
- An Analysis of Multilingual FActScoreVu Trong Kim, Michael Krumdick, Varshini Reddy, Franck Dernoncourt 等EMNLP 2024 · 被引用 2 次
- A Unified Encoder-Decoder Framework with Entity MemoryZhihan Zhang, Wenhao Yu, Chenguang Zhu, Meng JiangEMNLP 2022 · 被引用 9 次
- WebIE: Faithful and Robust Information Extraction on the WebChenxi Whitehouse, Clara Vania, Alham Fikri Aji, Christos Christodoulopoulos 等ACL 2023 · 被引用 3 次
- Wikiformer: Pre-training with Structured Information of Wikipedia for Ad-Hoc RetrievalWeihang Su, Qingyao Ai, Xiangsheng Li, Jia Chen 等AAAI 2024
