Monolingual Transfer Learning via Bilingual Translators for Style-Sensitive Paraphrase Generation
Tomoyuki Kajiwara, Biwa Miura, Yuki Arase
Abstract
We tackle the low-resource problem in style transfer by employing transfer learning that utilizes abundantly available raw corpora. Our method consists of two steps: pre-training learns to generate a semantically equivalent sentence with an input assured grammaticality, and fine-tuning learns to add a desired style. Pre-training has two options, auto-encoding and machine translation based methods. Pre-training based on AutoEncoder is a simple way to learn these from a raw corpus. If machine translators are available, the model can learn more diverse paraphrasing via roundtrip translation. After these, fine-tuning achieves high-quality paraphrase generation even in situations where only 1k sentence pairs of the parallel corpus for style transfer is available. Experimental results of formality style transfer indicated the effectiveness of both pre-training methods and the method based on roundtrip translation achieves state-of-the-art performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 38c55db3-8a9d-45d1-8a3b-b9a779f73072Related papers
- Few-shot Controllable Style Transfer for Low-Resource Multilingual SettingsKalpesh Krishna, Deepak Nathani, Xavier Garcia, Bidisha Samanta et al.ACL 2022 · 28 citations
- Reformulating Unsupervised Style Transfer as Paraphrase GenerationKalpesh Krishna, John Wieting, Mohit IyyerEMNLP 2020 · 9 citations
- Generic resources are what you need: Style transfer tasks without task-specific parallel training dataHuiyuan Lai, Antonio Toral, Malvina NissimEMNLP 2021 · 14 citations
- Paraphrasing as Zero-shot Translation with Feature-guided Diversity EnhancementZiyue Yan, Hongying Zan, Xinglin Lyu, Hongfei XuACL 2026
- A Dataset for Low-Resource Stylized Sequence-to-Sequence GenerationYu Wu, Yunli Wang, Shujie LiuAAAI 2020 · 24 citations
