Improving Summarization with Human Edits
Zonghai Yao, Benjamin J. Schloss, Sai P. Selvaraj
摘要
Recent work has shown the promise of learning with human feedback paradigms to produce human-determined high-quality text. Existing works use human feedback to train large language models (LLMs) in general domain abstractive summarization and have obtained summary quality exceeding traditional likelihood training. In this paper, we focus on a less explored form of human feedback – Human Edits. We propose Sequence Alignment (un)Likelihood Training (SALT), a novel technique to use both the human-edited and model-generated data together in the training loop. In addition, we demonstrate simulating Human Edits with ground truth summaries coming from existing training data – Imitation edits, along with the model-generated summaries obtained after the training, to reduce the need for expensive human-edit data. In our experiments, we extend human feedback exploration from general domain summarization to medical domain summarization. Our results demonstrate the effectiveness of SALT in improving the summary quality with Human and Imitation Edits. Through additional experiments, we show that SALT outperforms the conventional RLHF method (designed for human preferences) – DPO, when applied to human-edit data. We hope the evidence in our paper prompts researchers to explore, collect, and better use different human feedback approaches scalably.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Can AI writing be salvaged? Mitigating Idiosyncrasies and Improving Human-AI Alignment in the Writing Process through EditsTuhin Chakrabarty, Philippe Laban, Chien-Sheng WuCHI 2025 · 被引用 14 次
- SYNFAC-EDIT: Synthetic Imitation Edit Feedback for Factual Alignment in Clinical SummarizationPrakamya Mishra, Zonghai Yao, Parth Vashisht, Feiyun Ouyang 等EMNLP 2024 · 被引用 5 次
- LLM-Based Multi-Agent Systems for Clinical Workflows: A Survey of AI HospitalsZonghai Yao, Hong YuACL 2026
- Is the Top Still Spinning? Evaluating Subjectivity in Narrative UnderstandingMelanie Subbiah, Akankshya Mishra, Grace Kim, Liyan Tang 等EMNLP 2025
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach 等ICLR 2022 · 被引用 1,976 次
- The Flan Collection: Designing Data and Methods for Effective Instruction TuningShayne Longpre, Le Hou, Tu Vu, Albert Webson 等ICML 2023 · 被引用 908 次
相关 Paper
- Nash Learning from Human FeedbackRémi Munos, Michal Valko, Daniele Calandriello, Mohammad Gheshlaghi Azar 等ICML 2024 · 被引用 212 次
- Unsupervised Opinion Summarization with Noising and DenoisingReinald Kim Amplayo, Mirella LapataACL 2020 · 被引用 8 次
- An Imitation Learning Curriculum for Text Editing with Non-Autoregressive ModelsSweta Agrawal, Marine CarpuatACL 2022
- Teaching Language Models to Hallucinate Less with Synthetic TasksErik Jones, Hamid Palangi, Clarisse Simões, Varun Chandrasekaran 等ICLR 2024 · 被引用 43 次
- On Improving Summarization Factual Consistency from Natural Language FeedbackYixin Liu, Budhaditya Deb, Milagro Teruel, Aaron Halfaker 等ACL 2023 · 被引用 10 次
