Modelling Sentence Pairs via Reinforcement Learning: An Actor-Critic Approach to Learn the Irrelevant Words
Mahtab Ahmed, Robert E. Mercer
Abstract
Learning sentence representation is a fundamental task in Natural Language Processing. Most of the existing sentence pair modelling architectures focus only on extracting and using the rich sentence pair features. The drawback of utilizing all of these features makes the learning process much harder. In this study, we propose a reinforcement learning (RL) method to learn a sentence pair representation when performing tasks like semantic similarity, paraphrase identification, and question-answer pair modelling. We formulate this learning problem as a sequential decision making task where the decision made in the current state will have a strong impact on the following decisions. We address this decision making with a policy gradient RL method which chooses the irrelevant words to delete by looking at the sub-optimal representation of the sentences being compared. With this policy, extensive experiments show that our model achieves on par performance when learning task-specific representations of sentence pairs without needing any further knowledge like parse trees. We suggest that the simplicity of each task inference provided by our RL model makes it easier to explain.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Related papers
- Efficient Unsupervised Sentence Compression by Fine-tuning Transformers with Reinforcement LearningDemian Gholipour Ghalandari, Chris Hokamp, Georgiana IfrimACL 2022
- Teaching LLMs Human-Like Editing of Inappropriate Argumentation via Reinforcement LearningTimon Ziegenbein, Maja Stahl, Henning WachsmuthACL 2026
- Reinforcement Learning Based Graph-to-Sequence Model for Natural Question GenerationYu Chen, Lingfei Wu, Mohammed J. ZakiICLR 2020 · 167 citations
- Unsupervised Paraphrasing via Deep Reinforcement LearningA. B. Siddique, Samet Oymak, Vagelis HristidisKDD 2020 · 27 citations
- Enhancing Reinforcement Learning with Label-Sensitive Reward for Natural Language UnderstandingKuo Liao, Shuang Li, Meng Zhao, Liqun Liu et al.ACL 2024 · 2 citations
