BiSECT: Learning to Split and Rephrase Sentences with Bitexts
Joongwon Kim, Mounica Maddela, Reno Kriz, Wei Xu, Chris Callison-Burch
摘要
An important task in NLP applications such as sentence simplification is the ability to take a long, complex sentence and split it into shorter sentences, rephrasing as necessary. We introduce a novel dataset and a new model for this 'split and rephrase' task. Our BISECT training data consists of 1 million long English sentences paired with shorter, meaning-equivalent English sentences. We obtain these by extracting 1-2 sentence alignments in bilingual parallel corpora and then using machine translation to convert both sides of the corpus into the same language. BISECT contains higher quality training examples than previous Split and Rephrase corpora, with sentence splits that require more significant modifications. We categorize examples in our corpus, and use these categories in a novel model that allows us to target specific regions of the input sentence to be split and edited. Moreover, we show that models trained on BISECT can perform a wider variety of split operations and improve upon previous state-of-the-art approaches in automatic and human evaluations. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- LENS: A Learnable Evaluation Metric for Text SimplificationMounica Maddela, Yao Dou, David Heineman, Wei XuACL 2023 · 被引用 21 次
- Multilingual Simplification of Medical TextsSebastian Joseph, Kathryn Kazanas, Keziah Reina, Vishnesh J. Ramanathan 等EMNLP 2023 · 被引用 16 次
- Improving Large-scale Paraphrase Acquisition and GenerationYao Dou, Chao Jiang, Wei XuEMNLP 2022 · 被引用 11 次
- DEplain: A German Parallel Corpus with Intralingual Translations into Plain Language for Sentence and Document SimplificationRegina Stodden, Omar Momen, Laura KallmeyerACL 2023 · 被引用 3 次
- Aya Model: An Instruction Finetuned Open-Access Multilingual Language ModelAhmet Üstün, Viraat Aryabumi, Zheng Xin Yong, Wei-Yin Ko 等ACL 2024
它引用的顶会 Paper1
相关 Paper
- ASSET: A Dataset for Tuning and Evaluation of Sentence Simplification Models with Multiple Rewriting TransformationsFernando Alva-Manchego, Louis Martin, Antoine Bordes, Carolina Scarton 等ACL 2020 · 被引用 12 次
- Neural CRF Model for Sentence Alignment in Text SimplificationChao Jiang, Mounica Maddela, Wuwei Lan, Yang Zhong 等ACL 2020 · 被引用 103 次
- SWiPE: A Dataset for Document-Level Simplification of Wikipedia PagesPhilippe Laban, Jesse Vig, Wojciech Kryscinski, Shafiq Joty 等ACL 2023 · 被引用 5 次
- Discourse Level Factors for Sentence Deletion in Text SimplificationYang Zhong, Chao Jiang, Wei Xu, Junyi Jessy LiAAAI 2020 · 被引用 57 次
- Evaluating LLMs for Portuguese Sentence Simplification with Linguistic InsightsArthur Mariano Rocha De Azevedo Scalercio, Elvis A. de Souza, Maria José Bocorny Finatto, Aline PaesACL 2025 · 被引用 2 次
