Multi-Dimensional Optimization for Text Summarization via Reinforcement Learning
Sangwon Ryu, Heejin Do, Yunsu Kim, Gary Lee, Jungseul Ok
摘要
The evaluation of summary quality encompasses diverse dimensions such as consistency, coherence, relevance, and fluency. However, existing summarization methods often target a specific dimension, facing challenges in generating well-balanced summaries across multiple dimensions. In this paper, we propose multiobjective reinforcement learning tailored to generate balanced summaries across all four dimensions. We introduce two multi-dimensional optimization (MDO) strategies for adaptive learning: 1) MDO min , rewarding the current lowest dimension score, and 2) MDO pro , optimizing multiple dimensions similar to multi-task learning, resolves conflicting gradients across dimensions through gradient projection. Unlike prior ROUGE-based rewards relying on reference summaries, we use a QA-based reward model that aligns with human preferences. Further, we discover the capability to regulate the length of summaries by adjusting the discount factor, seeking the generation of concise yet informative summaries that encapsulate crucial points. Our approach achieved substantial performance gains compared to baseline models on representative summarization datasets, particularly in the overlooked dimensions. * Equal contribution That this Act may be cited as the ``Federal Forage Fee Act of 1993''. SECTION 1. FINDINGS. (a) Findings.--Congress finds and declares that--(1) it is in the national interest that the public lands are producing and continue to produce water and soil conservation benefits, livestock forage, wildlife forage and recreation and other multiple use opportunities; (2) rangelands will continue to be … The results of the updated survey shall be incorporated into the calculation of the Non Fee Cost Differential as they become available.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Autoregressive Multi-trait Essay Scoring via Reinforcement Learning with Scoring-aware Multiple RewardsHeejin Do, Sangwon Ryu, Gary Geunbae LeeEMNLP 2024 · 被引用 5 次
- SPARD: Self-Paced Curriculum for RL Alignment via Integrating Reward Dynamics and Data UtilityXuyang Zhi, Peilun Zhou, Chengqiang Lu, Hang Lv 等ACL 2026 · 被引用 3 次
- Adaptive Planning for Multi-Attribute Controllable Summarization with Monte Carlo Tree SearchSangwon Ryu, Heejin Do, Yunsu Kim, Gary Geunbae Lee 等ACL 2026 · 被引用 2 次
- RESTL: Reinforcement Learning Guided by Multi-Aspect Rewards for Signal Temporal Logic TransformationYue Fang, Zhi Jin, Jie An, Hongshen Chen 等AAAI 2026 · 被引用 1 次
- Incorporating Self-Rewriting into Large Language Model Reasoning ReinforcementJiashu Yao, Heyan Huang, Shuang Zeng, Chuwei Luo 等AAAI 2026
它引用的顶会 Paper16
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 被引用 2,453 次
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine 等NeurIPS 2020 · 被引用 2,261 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang 等EMNLP 2023 · 被引用 549 次
相关 Paper
- A Multi-Document Coverage Reward for RELAXed Multi-Document SummarizationJacob Parnell, Inigo Jauregi Unanue, Massimo PiccardiACL 2022 · 被引用 16 次
- Radiology Report Generation via Multi-objective Preference OptimizationTing Xiao, Lei Shi, Peng Liu, Zhe Wang 等AAAI 2025 · 被引用 21 次
- Multi-document Summarization with Maximal Marginal Relevance-guided Reinforcement LearningYuning Mao, Yanru Qu, Yiqing Xie, Xiang Ren 等EMNLP 2020 · 被引用 43 次
- Confronting Reward Model Overoptimization with Constrained RLHFTed Moskovitz, Aaditya K. Singh, DJ Strouse, Tuomas Sandholm 等ICLR 2024 · 被引用 89 次
- No Reader Left Behind: Multi-Agent Summaries Everyone Can UnderstandJimin Jung, MyoungJin Kim, Jaehyung Seo, Heuiseok LimACL 2026
