Controlled Text Reduction
Aviv Slobodkin, Paul Roit, Eran Hirsch, Ori Ernst, Ido Dagan
摘要
Producing a reduced version of a source text, as in generic or focused summarization, inherently involves two distinct subtasks: deciding on targeted content and generating a coherent text conveying it. While some popular approaches address summarization as a single end-to-end task, prominent works support decomposed modeling for individual subtasks. Further, semi-automated text reduction is also very appealing, where users may identify targeted content while models would generate a corresponding coherent summary. In this paper, we focus on the second subtask, of generating coherent text given pre-selected content. Concretely, we formalize Controlled Text Reduction as a standalone task, whose input is a source text with marked spans of targeted content ("highlighting"). A model then needs to generate a coherent text that includes all and only the target information. We advocate the potential of such models, both for modular fully-automatic summarization, as well as for semi-automated human-in-the-loop use cases. Facilitating proper research, we crowdsource high-quality dev and test datasets for the task. Further, we automatically generate a larger "silver" training dataset from available summarization benchmarks, leveraging a pretrained summary-source alignment model. Finally, employing these datasets, we present a supervised baseline model, showing promising results and insightful analyses. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- : Evaluating Factual Consistency in Knowledge-Grounded Dialogues via Question Generation and Question AnsweringOr Honovich, Leshem Choshen, Roee Aharoni, Ella Neeman 等EMNLP 2021 · 被引用 101 次
- Bridging the Structural Gap Between Encoding and Decoding for Data-To-Text GenerationChao Zhao, Marilyn A. Walker, Snigdha ChaturvediACL 2020 · 被引用 82 次
- Coarse-to-Fine Query Focused Multi-Document SummarizationYumo Xu, Mirella LapataEMNLP 2020 · 被引用 76 次
- CTRLsum: Towards Generic Controllable Text SummarizationJunxian He, Wojciech Kryscinski, Bryan McCann, Nazneen Rajani 等EMNLP 2022 · 被引用 59 次
相关 Paper
- CoCon: A Self-Supervised Approach for Controlled Text GenerationAlvin Chan, Yew-Soon Ong, Bill Pung, Aston Zhang 等ICLR 2021 · 被引用 16 次
- Socratic Pretraining: Question-Driven Pretraining for Controllable SummarizationArtidoro Pagnoni, Alexander R. Fabbri, Wojciech Kryscinski, Chien-Sheng WuACL 2023 · 被引用 4 次
- EntSUM: A Data Set for Entity-Centric Extractive SummarizationMounica Maddela, Mayank Kulkarni, Daniel Preotiuc-PietroACL 2022 · 被引用 2 次
- Plan ahead: Self-Supervised Text Planning for Paragraph Completion TaskDongyeop Kang, Eduard H. HovyEMNLP 2020 · 被引用 13 次
- Length Control in Abstractive Summarization by Pretraining Information SelectionYizhu Liu, Qi Jia, Kenny Q. ZhuACL 2022 · 被引用 39 次
