SWiPE: A Dataset for Document-Level Simplification of Wikipedia Pages
Philippe Laban, Jesse Vig, Wojciech Kryscinski, Shafiq Joty, Caiming Xiong, Chien-Sheng Wu
摘要
Text simplification research has mostly focused on sentence-level simplification, even though many desirable edits-such as adding relevant background information or reordering contentmay require document-level context. Prior work has also predominantly framed simplification as a single-step, input-to-output task, only implicitly modeling the fine-grained, span-level edits that elucidate the simplification process. To address both gaps, we introduce the SWIPE dataset, which reconstructs the document-level editing process from English Wikipedia (EW) articles to paired Simple Wikipedia (SEW) articles. In contrast to prior work, SWIPE leverages the entire revision history when pairing pages in order to better identify simplification edits. We work with Wikipedia editors to annotate 5,000 EW-SEW document pairs, labeling more than 40,000 edits with proposed 19 categories. To scale our efforts, we propose several models to automatically label edits, achieving an F-1 score of up to 70.6, indicating that this is a tractable but challenging NLU task. Finally, we categorize the edits produced by several simplification models and find that SWIPE-trained models generate more complex edits while reducing unwanted edits.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Beyond the Chat: Executable and Verifiable Text-Editing with LLMsPhilippe Laban, Jesse Vig, Marti A. Hearst, Caiming Xiong 等UIST 2024 · 被引用 29 次
- Can AI writing be salvaged? Mitigating Idiosyncrasies and Improving Human-AI Alignment in the Writing Process through EditsTuhin Chakrabarty, Philippe Laban, Chien-Sheng WuCHI 2025 · 被引用 14 次
- Know Your Audience: The benefits and pitfalls of generating plain language summaries beyond the "general" audienceTal August, Kyle Lo, Noah A. Smith, Katharina ReineckeCHI 2024 · 被引用 11 次
- An Open Multilingual System for Scoring Readability of WikipediaMykola Trokhymovych, Indira Sen, Martin GerlachACL 2024
- InfoLossQA: Characterizing and Recovering Information Loss in Text SimplificationJan Trienes, Sebastian Joseph, Jörg Schlötterer, Christin Seifert 等ACL 2024
它引用的顶会 Paper9
- Neural CRF Model for Sentence Alignment in Text SimplificationChao Jiang, Mounica Maddela, Wuwei Lan, Yang Zhong 等ACL 2020 · 被引用 103 次
- Evaluating the Factual Consistency of Abstractive Text SummarizationWojciech Kryscinski, Bryan McCann, Caiming Xiong, Richard SocherEMNLP 2020 · 被引用 67 次
- On Faithfulness and Factuality in Abstractive SummarizationJoshua Maynez, Shashi Narayan, Bernd Bohnet, Ryan T. McDonaldACL 2020 · 被引用 54 次
- Document-Level Text Simplification: Dataset, Criteria and BaselineRenliang Sun, Hanqi Jin, Xiaojun WanEMNLP 2021 · 被引用 33 次
- arXivEdits: Understanding the Human Revision Process in Scientific WritingChao Jiang, Wei Xu, Samuel StevensEMNLP 2022 · 被引用 10 次
相关 Paper
- SIMSUM: Document-level Text Simplification via Simultaneous SummarizationSofia Blinova, Xinyu Zhou, Martin Jaggi, Carsten Eickhoff 等ACL 2023 · 被引用 11 次
- ASSET: A Dataset for Tuning and Evaluation of Sentence Simplification Models with Multiple Rewriting TransformationsFernando Alva-Manchego, Louis Martin, Antoine Bordes, Carolina Scarton 等ACL 2020 · 被引用 12 次
- Discourse Level Factors for Sentence Deletion in Text SimplificationYang Zhong, Chao Jiang, Wei Xu, Junyi Jessy LiAAAI 2020 · 被引用 57 次
- Cross-lingual Science Journalism: Select, Simplify and Rewrite Summaries for Non-expert ReadersMehwish Fatima, Michael StrubeACL 2023 · 被引用 2 次
- BiSECT: Learning to Split and Rephrase Sentences with BitextsJoongwon Kim, Mounica Maddela, Reno Kriz, Wei Xu 等EMNLP 2021 · 被引用 14 次
