Style Equalization: Unsupervised Learning of Controllable Generative Sequence Models
Jen-Hao Rick Chang, Ashish Shrivastava, Hema Koppula, Xiaoshuai Zhang, Oncel Tuzel
Abstract
Controllable generative sequence models with the capability to extract and replicate the style of specific examples enable many applications, including narrating audiobooks in different voices, auto-completing and auto-correcting written handwriting, and generating missing training samples for downstream recognition tasks. However, under an unsupervised-style setting, typical training algorithms for controllable sequence generative models suffer from the training-inference mismatch, where the same sample is used as content and style input during training but unpaired samples are given during inference. In this paper, we tackle the training-inference mismatch encountered during unsupervised learning of controllable generative sequence models. The proposed method is simple yet effective, where we use a style transformation module to transfer target style information into an unrelated style input. This method enables training using unpaired content and style samples and thereby mitigate the training-inference mismatch. We apply style equalization to text-to-speech and text-to-handwriting synthesis on three datasets. We conduct thorough evaluation, including both quantitative and qualitative user studies. Our results show that by mitigating the training-inference mismatch with the proposed style equalization, we achieve style replication scores comparable to real data in our user studies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ea849db0-add2-4455-b2bc-97298bf28589Builds on14
- Alias-Free Generative Adversarial NetworksTero Karras, Miika Aittala, Samuli Laine, Erik Härkönen et al.NeurIPS 2021 · 2,126 citations
- Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment SearchJaehyeon Kim, Sungwon Kim, Jungil Kong, Sungroh YoonNeurIPS 2020 · 663 citations
- FastSpeech 2: Fast and High-Quality End-to-End Text to SpeechYi Ren, Chenxu Hu, Xu Tan, Tao Qin et al.ICLR 2021 · 513 citations
- How much Position Information Do Convolutional Neural Networks Encode?Md. Amirul Islam, Sen Jia, Neil D. B. BruceICLR 2020 · 392 citations
- Unsupervised Speech Decomposition via Triple Information BottleneckKaizhi Qian, Yang Zhang, Shiyu Chang, Mark Hasegawa-Johnson et al.ICML 2020 · 210 citations
Related papers
- Exploring Contextual Word-level Style Relevance for Unsupervised Style TransferChulun Zhou, Liangyu Chen, Jiachen Liu, Xinyan Xiao et al.ACL 2020 · 34 citations
- Autoregressive Stylized Motion Synthesis With Generative FlowYu-Hui Wen, Zhipeng Yang, Hongbo Fu, Lin Gao et al.CVPR 2021
- Improving Zero-Shot Voice Style Transfer via Disentangled Representation LearningSiyang Yuan, Pengyu Cheng, Ruiyi Zhang, Weituo Hao et al.ICLR 2021 · 64 citations
- Plug and Play Autoencoders for Conditional Text GenerationFlorian Mai, Nikolaos Pappas, Ivan Montero, Noah A. Smith et al.EMNLP 2020 · 24 citations
- Collaborative Learning of Bidirectional Decoders for Unsupervised Text Style TransferYun Ma, Yangbin Chen, Xudong Mao, Qing LiEMNLP 2021 · 6 citations
