Sequence Generation with Optimal-Transport-Enhanced Reinforcement Learning
Liqun Chen, Ke Bai, Chenyang Tao, Yizhe Zhang, Guoyin Wang, Wenlin Wang, Ricardo Henao, Lawrence Carin
摘要
Reinforcement learning (RL) has been widely used to aid training in language generation. This is achieved by enhancing standard maximum likelihood objectives with user-specified reward functions that encourage global semantic consistency. We propose a principled approach to address the difficulties associated with RL-based solutions, namely, high-variance gradients, uninformative rewards and brittle training. By leveraging the optimal transport distance, we introduce a regularizer that significantly alleviates the above issues. Our formulation emphasizes the preservation of semantic features, enabling end-to-end training instead of ad-hoc fine-tuning, and when combined with RL, it controls the exploration space for more efficient model updates. To validate the effectiveness of the proposed solution, we perform a comprehensive evaluation covering a wide variety of NLP tasks: machine translation, abstractive text summarization and image caption, with consistent improvements over competing solutions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Re-evaluating Word Mover's DistanceRyoma Sato, Makoto Yamada, Hisashi KashimaICML 2022 · 被引用 25 次
- Non-Parallel Text Style Transfer with Self-Parallel SupervisionRuibo Liu, Chongyang Gao, Chenyan Jia, Guangxuan Xu 等ICLR 2022 · 被引用 19 次
- Adaptive Prior-Dependent Correction Enhanced Reinforcement Learning for Natural Language GenerationWei Cheng, Ziyan Luo, Qiyue YinAAAI 2021 · 被引用 1 次
- AnomSeer: Reinforcing Multimodal LLMs to Reason for Time-Series Anomaly DetectionJunru Zhang, Lang Feng, Haoran Shi, Xu Guo 等ICML 2026
相关 Paper
- Semantic-aware Wasserstein Policy Regularization for Large Language Model AlignmentByeonghu Na, Hyungho Na, Yeongmin Kim, Suhyeon Jo 等ICLR 2026 · 被引用 2 次
- Optimal Transport for LLM Reward Modeling from Noisy FeedbackLicheng Pan, Haocheng Yang, Haoxuan Li, Yunsheng Lu 等ICML 2026
- Improving Text Generation with Student-Forcing Optimal TransportJianqiao Li, Chunyuan Li, Guoyin Wang, Hao Fu 等EMNLP 2020 · 被引用 11 次
- Factually Consistent Summarization via Reinforcement Learning with Textual Entailment FeedbackPaul Roit, Johan Ferret, Lior Shani, Roee Aharoni 等ACL 2023 · 被引用 21 次
- Pose-RFT: Aligning MLLMs for 3D Pose Generation via Hybrid Action Reinforcement Fine-TuningBao Li, Xiaomei Zhang, Miao Xu, Zhaoxin Fan 等ICLR 2026
