Controlled Text Reduction
Aviv Slobodkin, Paul Roit, Eran Hirsch, Ori Ernst, Ido Dagan
Abstract
Producing a reduced version of a source text, as in generic or focused summarization, inherently involves two distinct subtasks: deciding on targeted content and generating a coherent text conveying it. While some popular approaches address summarization as a single end-to-end task, prominent works support decomposed modeling for individual subtasks. Further, semi-automated text reduction is also very appealing, where users may identify targeted content while models would generate a corresponding coherent summary. In this paper, we focus on the second subtask, of generating coherent text given pre-selected content. Concretely, we formalize Controlled Text Reduction as a standalone task, whose input is a source text with marked spans of targeted content ("highlighting"). A model then needs to generate a coherent text that includes all and only the target information. We advocate the potential of such models, both for modular fully-automatic summarization, as well as for semi-automated human-in-the-loop use cases. Facilitating proper research, we crowdsource high-quality dev and test datasets for the task. Further, we automatically generate a larger "silver" training dataset from available summarization benchmarks, leveraging a pretrained summary-source alignment model. Finally, employing these datasets, we present a supervised baseline model, showing promising results and insightful analyses. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 581247f2-243e-4f09-a819-ed24d1497b64Cited by top-tier papers1
Ask how each one uses itBuilds on6
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- : Evaluating Factual Consistency in Knowledge-Grounded Dialogues via Question Generation and Question AnsweringOr Honovich, Leshem Choshen, Roee Aharoni, Ella Neeman et al.EMNLP 2021 · 101 citations
- Bridging the Structural Gap Between Encoding and Decoding for Data-To-Text GenerationChao Zhao, Marilyn A. Walker, Snigdha ChaturvediACL 2020 · 82 citations
- Coarse-to-Fine Query Focused Multi-Document SummarizationYumo Xu, Mirella LapataEMNLP 2020 · 76 citations
- CTRLsum: Towards Generic Controllable Text SummarizationJunxian He, Wojciech Kryscinski, Bryan McCann, Nazneen Rajani et al.EMNLP 2022 · 59 citations
Related papers
- CoCon: A Self-Supervised Approach for Controlled Text GenerationAlvin Chan, Yew-Soon Ong, Bill Pung, Aston Zhang et al.ICLR 2021 · 16 citations
- Socratic Pretraining: Question-Driven Pretraining for Controllable SummarizationArtidoro Pagnoni, Alexander R. Fabbri, Wojciech Kryscinski, Chien-Sheng WuACL 2023 · 4 citations
- EntSUM: A Data Set for Entity-Centric Extractive SummarizationMounica Maddela, Mayank Kulkarni, Daniel Preotiuc-PietroACL 2022 · 2 citations
- Plan ahead: Self-Supervised Text Planning for Paragraph Completion TaskDongyeop Kang, Eduard H. HovyEMNLP 2020 · 13 citations
- Length Control in Abstractive Summarization by Pretraining Information SelectionYizhu Liu, Qi Jia, Kenny Q. ZhuACL 2022 · 39 citations
