Generating Summaries with Controllable Readability Levels
Leonardo F. R. Ribeiro, Mohit Bansal, Markus Dreyer
摘要
Readability refers to how easily a reader can understand a written text. Several factors affect the readability level, such as the complexity of the text, its subject matter, and the reader's background knowledge. Generating summaries based on different readability levels is critical for enabling knowledge consumption by diverse audiences. However, current text generation approaches lack refined control, resulting in texts that are not customized to readers' proficiency levels. In this work, we bridge this gap and study techniques to generate summaries at specified readability levels. Unlike previous methods that focus on a specific readability level (e.g., lay summarization), we generate summaries with fine-grained control over their readability. We develop three text generation techniques for controlling readability: (1) instruction-based readability control, (2) reinforcement learning to minimize the gap between requested and observed readability and (3) a decoding approach that uses lookahead to estimate the readability of upcoming decoding steps. We show that our generation methods significantly improve readability control on news summarization (CNN/DM dataset), as measured by various readability metrics and human judgement, establishing strong baselines for controllable readability in summarization. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Autoregressive Multi-trait Essay Scoring via Reinforcement Learning with Scoring-aware Multiple RewardsHeejin Do, Sangwon Ryu, Gary Geunbae LeeEMNLP 2024 · 被引用 5 次
- Standardize: Aligning Language Models with Expert-Defined Standards for Content GenerationJoseph Marvin Imperial, Gail Forey, Harish Tayyar MadabushiEMNLP 2024 · 被引用 3 次
- RSA-Control: A Pragmatics-Grounded Lightweight Controllable Text Generation FrameworkYifan Wang, Vera DembergEMNLP 2024 · 被引用 2 次
- LSRP: A Leader-Subordinate Retrieval Framework for Privacy-Preserving Cloud-Device CollaborationYingyi Zhang, Pengyue Jia, Xianneng Li, Derong Xu 等KDD 2025 · 被引用 2 次
- Introducing Spotlight: A Novel Approach for Generating Captivating Key Information from DocumentsAnkan Mullick, Sombit Bose, Rounak Saha, Ayan Kumar Bhowmick 等EMNLP 2025
它引用的顶会 Paper16
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 被引用 2,453 次
- Towards a Unified Multi-Dimensional Evaluator for Text GenerationMing Zhong, Yang Liu, Da Yin, Yuning Mao 等EMNLP 2022 · 被引用 103 次
- Automated Lay Language Summarization of Biomedical Scientific ReviewsYue Guo, Wei Qiu, Yizhong Wang, Trevor CohenAAAI 2021 · 被引用 100 次
相关 Paper
- Controlling Pre-trained Language Models for Grade-Specific Text SimplificationSweta Agrawal, Marine CarpuatEMNLP 2023 · 被引用 5 次
- Length Control in Abstractive Summarization by Pretraining Information SelectionYizhu Liu, Qi Jia, Kenny Q. ZhuACL 2022 · 被引用 39 次
- CTRLsum: Towards Generic Controllable Text SummarizationJunxian He, Wojciech Kryscinski, Bryan McCann, Nazneen Rajani 等EMNLP 2022 · 被引用 59 次
- Text Generation by Learning from DemonstrationsRichard Yuanzhe Pang, He HeICLR 2021 · 被引用 88 次
- CEFR-Based Sentence Difficulty Annotation and AssessmentYuki Arase, Satoru Uchida, Tomoyuki KajiwaraEMNLP 2022 · 被引用 17 次
